Processing image, along with new camera pose and / or new timestamp, to generate predicted image from the new camera pose and / or at the new timestamp
An image generation network processes base images with new camera poses and timestamps to generate predicted images, addressing the challenge of synthesizing complex scenes with arbitrary poses and dynamics, enhancing robotic control and reducing resource usage.
Patent Information
- Application Number
- PCT/US2025/030415
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-22
- Filing Date
- 2025-05-21
- Publication Date
- 2025-11-27
AI Technical Summary
Existing image generation techniques struggle with synthesizing novel views of complex scenes involving arbitrary camera poses and dynamic elements, often compromising three-dimensional scene representation and consistency, especially when dealing with static scenes and out-of-distribution inputs.
An image generation network is trained and utilized to process base images with new camera poses and timestamps, enabling the generation of predicted images that reflect scenes from new angles and times, incorporating multi-guidance training to handle varied data sets, including those lacking full pose or timestamp information.
The approach allows for the generation of novel views of three-dimensional scenes with temporal dynamics from limited input data, improving the ability to synthesize views from arbitrary camera poses and at different points in time, and reducing resource usage and wear on robotic systems by generating synthetic training images.
Smart Images

Figure US2025030415_27112025_PF_FP_ABST
Abstract
Description
PROCESSING IMAGE, ALONG WITH NEW CAME A POSE AND / OR NEW TIMESTAMP, TO GENERATEPREDICTED IMAGE FROM THE NEW CAMERA POSE AND / OR AT THE NEW TIMESTAMPAtty Docket: DEEP-0020-WQ-01Background
[0001] Some approaches to image generation from limited input images often struggle with synthesizing novel views of complex scenes, particularly those involving arbitrary camera poses and dynamic elements that change over time. For example, some techniques are primarily designed for generating views of isolated objects against simple backgrounds or are restricted to limited camera trajectories.
[0002] For example, some techniques focus on generating new views by conditioning diffusion models on a single input image and a relative camera pose. While these approaches can generate plausible images, they are not well-suited for jointly modeling multiple views. As a result, the consistency of the generated three-dimensional scene representation can be compromised.
[0003] Further, certain diffusion models for novel view synthesis may incorporate attention mechanisms to leverage epipolar geometry and improve three-dimensional consistency. Other approaches have fine-tuned text-to-video models with temporal attention layers or attention layers limited to overlapping regions to implicitly model the joint distribution of generated views and their camera extrinsics. However, these models can exhibit several drawbacks. They may struggle with static scenes due to residual dynamics from the underlying video models, resulting in undesirable motion in generated static views. Three-dimensional inconsistencies and low fidelity can also persist. Additionally, these models may show poor generalization when applied to out-of-distribution image inputs.Summary
[0004] Implementations disclosed herein relate to training and / or using an image generation network for generating predicted image(s) from base image(s), such as a base image that is a real image captured by a real camera.
[0005] In generating the predicted image, some implementations process, using the image generation network and along with a base image, a new camera pose and a timestamp. For example, a positional encoding of the new camera pose and a relative timestamp (e.g., relative to a zero or other base time of the base image) can be processed along with the base image. The new camera pose differs from a camera pose of the base image and the timestamp reflects a time that is before or that is after a time of the base image. In those implementations, the predicted image captures the scene that is captured by the base image(s), but is from the processed camera pose and is at a time reflected by the processed timestamp. For example, being from the processed camera pose will cause the scene to be captured, in the predicted image, from a new angle and will potentially include additional object(s) predicted to be in the scene, despite the object(s) not being in the base image. As another example, being at the time reflected by the timestamp will cause, in the predicted image, dynamic object(s) in the scene to have moved (or otherwise changed) in accordance with the time. For instance, for a timestamp that indicates three seconds in the future, walking human(s) in the scene may have progressed a few meter(s) while moving car(s) in the scene may have progressed tens of meter(s).
[0006] Accordingly, some implementations disclosed herein process base image(s), a new camera pose, and a new timestamp to generate a predicted image that reflects the scene of the base image but is from the new camera pose and that is progressed (or regressed) to a time corresponding to the timestamp.
[0007] Predicted images generated according to implementations disclosed herein can be utilized for various purposes. Such purposes include generating synthetic training images based on limited real images and training robotic control policies and / or other control policies based on such synthetic training images. Such purposes additionally or alternatively include generating predicted images and utilizing those predicted images in active robotic control. For example, a limited quantity of real images of only parts of an environment can be captured utilizing a real camera of a real robot and techniques disclosed herein can be utilized to generate predicted images of the remaining parts of the environment. Such predicted images can be utilized in planning or other robotic control by the real robot, enabling the real robotcontrol policy / policies to interact with and / or avoid object(s) that may not be present in the real images and / or that may be in different earlier in time position(s) in the real images. Utilizing the predicted images, in lieu of the real robot navigating its base and / or independently adjusting the pose of the real camera, can lessen resource usage by the robot and / or can lessen wear and tear on robotic component(s). Further, the predicted image(s) can capture area(s) that cannot be captured by the real camera (e.g., due to mobility constraint(s) of the real robot) and can, when future timestamps are utilized, capture an environment in a state that has yet to occur.
[0008] In various implementations disclosed herein the same image generation network can be utilized to generate a predicted image that is from a new camera pose but not at a different time and / or to generate a predicted image that is at a different time but not from a new camera pose. For example, to generate a predicted image that is from a new camera pose but not at a different time some implementations process, using the image generation network and along with a base image, a new camera pose but without a timestamp. For instance, the timestamp may not be processed and instead a camera pose feature-wise linear modulation ( Fi LM ) layer, of the image generation network and that is used for processing camera poses, can be configured to behave like the identify function. As another example, to generate a predicted image that is from the same camera pose but at a different time some implementations process, using the image generation network and along with a base image, a new timestamp but without a new camera pose. For instance, the timestamp may not be processed and instead a timestamp FiLM layer, of the image generation network and that is used for processing timestamps, can be configured to behave like the identify function. In these and other manners, the same image generation network can be utilized to generate various types of predicted images including those that only have a new camera pose, those that only progress or regress time, and those that both have a new camera pose and that progress or regress time.
[0009] Some implementations that relate to training the image generation network utilize multi-guidance that enables varied types of data sets to be utilized during training and / or that enables the image generation network to condition to multiple elements of input including theinput image(s), new pose(s), and timestamp(s) - and optionally condition to subset(s) of the multiple element(s) (e.g., input image(s) and timestamp(s) without any conditioning to new pose(s)). The varied types of data sets can include, for example, pose-less instances, timestamp-less instances, and / or full instances. The pose-less instances include real values for the corresponding conditioning images and for the corresponding conditioning timestamps, but include null or default values for the corresponding conditioning camera poses. The timestamp-less instances that include real values for the corresponding conditioning images and for the corresponding camera poses, but include null or default values for the corresponding conditioning timestamps. The full instances include real values for the corresponding conditioning images, for the corresponding camera poses, and for the corresponding conditioning timestamps.
[0010] In implementations that utilize multi-guidance, iterations of training are performed based on the training instances, including no-drop iterations, camera pose drop iterations, and timestamp drop iterations. In the no-drop iterations, nothing is dropped from the corresponding training inputs for the training instances used in the no-drop iterations. In the camera pose drop iterations only the corresponding conditioning camera pose, for the training instances used in the camera pose drop iterations, is dropped from the corresponding training inputs. In the timestamp drop iterations only the corresponding conditioning timestamp, for the training instances used in the timestamp pose drop iterations, is dropped from the corresponding training inputs.
[0011] Optionally, pose-less instances and timestamp-less instances can be prevented from being utilized in the no-drop iterations, the timestamp-less instances can be prevented from being utilized in the no-drop iterations and the camera pose drop iterations, the pose-less instances can be prevented from being utilized in the no-drop iterations and the timestamp drop iterations, and / or the full instances can be caused to be used in the no-drop iterations. In these and other manners a full range of types of data sets can be utilized in robustly training the image generation network. This can prevent the need for computationally expensive generation of vast quantities of full instances and / or can enable utilization of timestamp-less and / or camera-less instances that can represent a more diverse set of scenes, such asuncalibrated outdoor scenes. Additionally or alternatively, and as referenced above, this can enable the trained image generation network to condition to multiple elements of input including the input image(s), new pose(s), and timestamp(s) - and optionally condition to subset(s) of the multiple element(s).
[0012] Some implementations disclosed herein includes systems and methods for generating predicted images of a scene based on a limited number of conditioning image(s), camera poses, and timestamps. The system can comprise an image generation network, such as an image diffusion model, configured to receive a conditioning image captured at a first time from a first camera pose. The image generation network is further configured to receive a second camera pose that differs from the first camera pose and a timestamp reflecting a second time that is different from the first time. The image generation network is configured to process the conditioning image, the second camera pose, and the timestamp to generate a predicted image that reflects the scene of the conditioning image but is rendered from the second camera pose and depicts the scene at the second time.
[0013] Some implementations of the present disclosure enable the generation of novel views of three-dimensional scenes with temporal dynamics from limited input data. Some additional or alternative implementations improve the ability to synthesize views from arbitrary camera poses and at different points in time.
[0014] Various implementations provide several advantages, including the ability to generate synthetic training data for applications such as robotic control, enabling robots to predict environmental states beyond their immediate sensor input or current time. The disclosed implementations also provide a method for training the image generation network using a variety of datasets, including those lacking full pose or timestamp information, through a multiguidance training approach.
[0015] In some implementations, the image generation network includes a plurality of chained feature-wise linear modulation (FiLM ) layers for processing the conditioning images, camera poses, and timestamps. Training can involve performing iterations where one or more of the conditioning inputs are intentionally dropped, allowing the network to learn to generate images based on subsets of the available input information. This training approach can utilizetraining instances with varying degrees of completeness, such as instances that include real values for conditioning images and timestamps but null values for camera poses, or instances with real values for conditioning images and camera poses but null values for timestamps. The training instances can also include full instances with real values for all conditioning inputs. The training process can include generating a loss based on comparing predicted outputs to ground truth images and updating the image generation network based on that loss.
[0016] After training, the image generation network can be used by computing devices to generate predicted images that reflect corresponding scenes of corresponding real images but are rendered from differing camera poses and / or depict the corresponding scenes at differing points in time. These predicted images can be utilized in various applications, including training robotic policies or directly controlling a robot whose camera captured the original real images. The camera pose and timestamp can be generated based on user interface inputs or by a robot itself.
[0017] The preceding is presented as a non-limiting overview of only some implementations disclosed herein. The appended paper and the claims provide additional details on those and other implementations.Brief Description of the Drawings
[0018] Figure 1 illustrates an example architecture of an image generation network.
[0019] Figure 2 illustrates an example method according to some implementations.
[0020] Figure 3 illustrates another example method according to some implementations.
[0021] Figure 4 illustrates an example robotic system according to some implementations.
[0022] Figure 5 illustrates an example computing system according to some implementations.
[0023] Figures 6A, 6B, 6C, and 6D illustrate example tables.Detailed Description
[0024] Prior to turning to the Figures, a non-limiting description of some example aspects of the disclosure is provided.
[0025] Implementations disclosed herein relate to training and / or using an image generation network for generating predicted image(s) from base image(s), such as a base image that is a real image captured by a real camera.
[0026] In generating the predicted image, some implementations process, using the image generation network and along with a base image, a new camera pose and a timestamp. For example, a positional encoding of the new camera pose and a relative timestamp (e.g., relative to a zero or other base time of the base image) can be processed along with the base image. The new camera pose differs from a camera pose of the base image and the timestamp reflects a time that is before or after a time of the base image. In those implementations the predicted image captures the scene that is captured by the base image(s), but is from the processed camera pose and is at a time reflected by the processed timestamp. For example, being from the processed camera pose will cause the scene to be captured, in the predicted image, from a new angle and will potentially include additional object(s) predicted to be in the scene, despite not being in the base image. As another example, being at the time reflected by the timestamp will cause, in the predicted image, dynamic object(s) in the scene to have moved (or otherwise changed) in accordance with the time. For instance, for a timestamp that indicates three seconds in the future walking human(s) in the scene may have progressed a few meter(s) while moving car(s) in the scene may have progressed tens of meter(s).
[0027] Accordingly, some implementations disclosed herein process base image(s), a new camera pose and a new timestamp to generate a predicted image that reflects the scene of the base image but is from the new camera pose and that is progressed (or regressed) to a time corresponding to the timestamp.
[0028] Predicted images generated according to implementations disclosed herein can be utilized for various purposes. Such purposes include generating synthetic training images based on limited real images and training robotic control policies and / or other control policies based on such synthetic training images. Such purposes additionally or alternatively include generating predicted images and utilizing those predicted images in active robotic control. For example, a limited quantity of real images of only parts of an environment can be captured utilizing a real camera of a real robot and techniques disclosed herein can be utilized to generate predicted images of the remaining parts of the environment. Such predicted images can be utilized in planning or other robotic control by the real robot, enabling the real robot control policy / policies to interact with and / or avoid object(s) that may not be present in thereal images and / or that may be in different earlier in time position(s) in the real images. Utilizing the predicted images, in lieu of the real robot navigating its base and / or independently adjusting the pose of the real camera, can lessen resource usage by the robot and / or lessen wear and tear on robotic component(s). Further, the predicted image(s) can capture area(s) that cannot be captured by the real camera (e.g., due to mobility constraint(s) of the real robot) and can, when future timestamps are utilized, capture an environment, in a state that has yet to occur.
[0029] In various implementations disclosed herein the same image generation network can be utilized to generate a predicted image that is from a new camera pose but not at a different time and / or to generate a predicted image that is at a different time but not from a new camera pose. For example, to generate a predicted image that is from a new camera pose but not at a different time some implementations process, using the image generation network and along with a base image, a new camera pose but without a timestamp. For instance, the timestamp may not be processed and instead a camera pose feature-wise linear modulation ( Fi LM ) layer, of the image generation network and that is used for processing camera poses, can be configured to behave like the identify function. As another example, to generate a predicted image that is from the same camera pose but at a different time some implementations process, using the image generation network and along with a base image, a new timestamp but without a new camera pose. For instance, the timestamp may not be processed and instead a timestamp FiLM layer, of the image generation network and that is used for processing timestamps, can be configured to behave like the identify function. In these and other manners, the same image generation network can be utilized to generate various types of predicted images including those that only have a new camera pose, those that only progress or regress time, and those that both have a new camera pose and that progress or regress time.
[0030] Some implementations that relate to training the image generation network utilize multi-guidance that enables varied types of data sets to be utilized during training and / or that enables the image generation network to condition to multiple elements of input including the input image(s), new pose(s), and timestamp(s) - and optionally condition to subset(s) of themultiple element(s) (e.g., input image(s) and timestamp(s) without any conditioning to new pose(s)). The varied types of data sets can include, for example, pose-less instances, timestamp-less instances, and / or full instances. The pose-less instances include real values for the corresponding conditioning images and for the corresponding conditioning timestamps, but include null or default values for the corresponding conditioning camera poses. The timestamp-less instances that include real values for the corresponding conditioning images and for the corresponding camera poses, but include null or default values for the corresponding conditioning timestamps. The full instances include real values for the corresponding conditioning images, for the corresponding camera poses, and for the corresponding conditioning timestamps.
[0031] In implementations that utilize multi-guidance, iterations of training are performed based on the training instances, including no-drop iterations, camera pose drop iterations, and timestamp drop iterations. In the no-drop iterations, nothing is dropped from the corresponding training inputs for the training instances used in the no-drop iterations. In the camera pose drop iterations only the corresponding conditioning camera pose, for the training instances used in the camera pose drop iterations, is dropped from the corresponding training inputs. In the timestamp drop iterations only the corresponding conditioning timestamp, for the training instances used in the timestamp pose drop iterations, is dropped from the corresponding training inputs.
[0032] Optionally, pose-less instances and timestamp-less instances can be prevented from being utilized in the no-drop iterations, the timestamp-less instances can be prevented from being utilized in the no-drop iterations and the camera pose drop iterations, the pose-less instances can be prevented from being utilized in the no-drop iterations and the timestamp drop iterations, and / or the full instances can be caused to be used in the no-drop iterations. In these and other manners a full range of types of data sets can be utilized in robustly training the image generation network. This can prevent the need for computationally expensive generation of vast quantities of full instances and / or can enable utilization of timestamp-less and / or camera-less instances that can represent a more diverse set of scenes, such as uncalibrated outdoor scenes. Additionally or alternatively, and as referenced above, this canenable the trained image generation network to condition to multiple elements of input including the input image(s), new pose(s), and timestamp(s) - and optionally condition to subset(s) of the multiple element(s).
[0033] Novel view synthesis (NVS) and three-dimensional generative models represent an advance in generative modeling, facilitating the synthesis of novel views of three-dimensional objects and scenes with control over camera pose. These models complement text-to- image / video models, with potential utility in generating three-dimensional assets, augmented reality systems, synthetic data for model training (for example, in robotics), and view interpolation processes. Some recent examples that utilize diffusion models do not generate explicit three-dimensional scenes. Rather, they leverage implicit three-dimensional knowledge for conditional view generation, operating as a form of image and pose conditioned image generation.
[0034] Recent work has attempted to train models for generic scenes and three-dimensional viewpoints, but with such recent work zero-shot performance on out-of-distribution scenes presents challenges. One difficulty lies in the scarcity of three-dimensional scene data. An additional difficulty is that existing data also relies on COLMAP for estimating camera pose, resulting in noisy poses with an indeterminate scale. This further complicates effective control in the single-image to three-dimensional task, as models sample from the distribution of plausible scales during inference.
[0035] Implementations disclosed herein extend NVS diffusion models in one or more aspects such as expanding from objects to scenes, accommodating free-form camera poses specified in meaningful physical units, and / or enabling simultaneous spatial and temporal control through camera pose and timestamp conditioning.
[0036] Implementations disclosed herein relate to utilizing a diffusion model, for NVS, conditioned on one or more images of arbitrary scenes, camera pose, and time. Such diffusion model is sometimes denoted herein as 4DiM. 4DiM can be trained on a combination of data sources, including posed three-dimensional images or video and unposed video, encompassing both indoor and outdoor scenes. This is made possible through various technical innovations that allow for training from data with missing information (for example, images without poseor time annotations) and sampling with distinct guidance weights applied to different conditioning images, poses, or times. Calibrated versions of video datasets posed via COLMAP are also utilized in various implementations. Calibrated data facilitates learning metric regularities in the world, such as typical sizes of common objects and spatial relationships, and permits specifying camera poses in meaningful physical units. As such, 4DiM generates three- dimensional consistent, multi-view images or video of dynamic scenes. Beyond sample generation, 4DiM can be utilized to facilitate a range of applications, including video-to-video translation, improved panoramic stitching, and / or training explicit three-dimensional models using score-distillation sampling.
[0037] 4DiM can be a pixel-based diffusion model for novel view synthesis conditioned on one or more images of arbitrary scenes, camera pose, and time. 4DiM can include a base model configured to generate multiple (e.g., 32 images) at a resolution (e.g., 64x64), and a multi-view super-resolution model configured to up-sample the multiple images (e.g., to a resolution of 256x256). 4DiM can be trained on a mixture of data sources, comprising posed and unposed video of indoor and outdoor scenes. 4DiM can be trained to enable zero-shot and fine-grained applications, including zero-shot video-to-video translation, improved panoramic stitching, and / or training explicit three-dimensional models with score-distillation sampling.
[0038] In contrast to the typical setting of Neural Radiance Fields (NeRF) where tens-to- hundreds of images are used as input for 3D reconstruction, pose-conditional diffusion models for NVS aim to extrapolate plausible, diverse, 3D consistent samples with as few as a single image input. Conditioning diffusion models on an image and relative camera pose was introduced as an effective alternative to prior few-view NVS methods, overcoming severe blur and floater artifacts . However, such models rely on neural architectures unsuitable for jointly modeling more than two views, and consequently require heuristics such as stochastic conditioning or Markovian sampling with a limited context window, for which it is hard to maintain 3D consistency.
[0039] Some work has proposed attention mechanisms leveraging epipolar geometry to improve the 3D consistency of image-to-image diffusion models for NVS, and more recently, fine-tuning text-to-video models with temporal attention layers or attention layers limited tooverlapping regions to model the joint distribution of generated views and their camera extrinsics. Such models however still suffer from several issues: difficulty with static scenes due to persistent dynamics from the underlying video models; persistent 3D inconsistencies and low fidelity; and poor generalization to out-of-distribution image inputs. Alternatives have been proposed to improve fidelity in multi-view diffusion models of 3D scenes, albeit sacrificing the ability to model free-form camera pose.
[0040] A concurrent trajectory of work on 3D extraction has recently emerged, where instead of directly training diffusion models for NVS, new techniques for sampling views parametrized as a NeRF with volume rendering are proposed such as Score Distillation Sampling (SDS) and Variational Score Distillation (VSD). With such work, a pre-existing diffusion model acts as a prior that drives the generation process. This enables, for example, text-to-3D using a text-to- image diffusion model. Using a pose-conditional diffusion model offers the unique advantage that the diffusion model can produce samples at the exact viewpoint specified for volume rendering during score distillation or NeRF postprocessing, which results in improved sample quality.
[0041] A continuous-time diffusion model can be utilized to learn the joint distribution over multiple views. This can be represented as the probability of generating images, given conditioning images, relative camera poses (extrinsics and intrinsics), and scalar, relative timestamps. For example, this can be represented as p(xc+1.N\x1:C, Pl: N,t1:iV), where xc+1.Nare generated images, x1;Care conditioning images, Pl: N are relative camera poses (extrinsics and intrinsics), and t1;Ware scalar, relative timestamps.
[0042] A loss function is selected that represents the error between the predicted and actual noise. The LI rather than L2 norm can be used, as it improves sample quality. The models can use the "v-parametrization", which helps stabilize training, and adopt the noise schedules. The current models process a number of images at a resolution of 256x256, where this number includes both conditioning and generated frames. To this end, the task is decomposed into two models. Images are first generated at an initial resolution, and then up-sampled using a superresolution model trained with noise conditioning augmentation.
[0043] While 3D assets, multi-view image data, and 4D data are limited, video data is available at scale and contains rich information about the 3D world, despite not having camera poses. One proposition involves training 4DiM on a large-scale dataset of (e.g., 30M) videos without pose annotations, jointly with 3D and 4D data. Video is shown to play a significant role in helping to regularize the model. The 3D datasets utilized to train 4DiM can include, for example, ScanNet++ and Matterport3D, which have metric scale, and more free-form camera poses compared to other common 3D datasets in the literature (e.g., CO3D and MVImgNet). Scenes from Street View can also be utilized, including posed panoramas with timestamps (i.e., it is a "4D" dataset). During training, views can be randomly sampled from Street View from the set of panorama images within a number of consecutive timesteps. Unposed videos can be sampled with a specified probability, and views from posed datasets are sampled otherwise. 3D datasets can be sampled in proportion to the number of scenes in each dataset.
[0044] One rich dataset for training 3D models is RealEstatelOK. It includes 10,000 video segments of static scenes for which SfM has been utilized to infer per-frame camera pose, but only up to an unknown length scale. The lack of metric scale makes training more difficult because metric regularities of the world are lost, and it becomes more difficult for users to specify the target camera poses or camera motions in any intuitive units. This can be problematic when conditioning on a single image, where scale itself otherwise becomes ambiguous. A calibrated version of RealEstatelOK can be created by regressing the unknown metric scale from a monocular depth estimation model.
[0045] Relatively little training data has both time and camera pose annotations. Rather, most 3D data represent static scenes, while video data rarely include camera pose. Finding a way to effectively condition on both camera pose and time in a way that allows for incomplete training data can be beneficial. A proposal involves chaining "Masked FiLM" layers for (positional encodings of) diffusion noise levels, per-pixel ray origins and directions, and video timestamps. When any of these conditioning signals is missing (due to incomplete training data or random dropout for classifier-free guidance), the FiLM layers are designed to reduce to the identity function, rather than simply setting missing values to zero. This avoids the networkfrom confusing a timestamp of zero with dropped or missing data. In practice, the FiLM shift can be replaced with zeros and the scale with ones.
[0046] Steering the model with the correct sampling hyperparameters can be beneficial, especially the guidance weights for classifier-free guidance (CFG). In its usual formulation, CFG treats all conditioning variables as "one big variable". In practice, placing a different weight on each variable can be important. For example, text requires high guidance weights, but high guidance weights on images can quickly lead to unwanted artifacts. Multi-guidance is proposed, where CFG is generalized to do exactly this, without making independence assumptions between conditioning variables. This begins from a classifier-guided formulation, where the aim is to sample with a number of conditioning signals. For example, where k conditioning signals,are to be sampled, this can be represented using the score:
[0047] V log
[0048] In the classifier-free formulation, this is equivalent to a combination of terms involving the marginal probability of the data and the conditional probability of the data given subsets of the conditioning variables. For example, this can be represented as
[0050] If the model is trained such that it drops out only a subset of the conditioning variables, or all of them, sampling can be performed with different guidance weights on each conditioning variable. The model can be trained to drop out conditioning signals with a specified probability, and it drops either timestamps, timestamps and poses, or all of timestamps, poses, and conditioning images. This can be a beneficial choice, since poses and timestamps without corresponding conditioning images do not convey useful information. As a non-limiting example, a guidance weight of 2.0 can be utilized on conditioning images, a weight of 4.0 on camera poses, and a weight of 1.0 for timestamps.
[0051] Evaluating 4D generative models is challenging. Typically, methods for NVS are evaluated using generation quality. In addition, for pose-conditioned generation it is desirable to find metrics that capture 3D consistency (rendered views should be grounded in a consistent 3D scene) and pose alignment (the camera should move as expected). Furthermore, evaluating temporal conditioning requires some measure that captures the motion of dynamic content.Several existing metrics are leveraged, and new ones are proposed to cover all these aspects. Each one is covered in detail herein.
[0052] FID is one metric for image quality, but if used in isolation it can be uninformative and a poor objective for hyper-parameter selection. For example, FID scores for text-to-image models are optimal with low CFG weights, but text-image alignment suffers under this setting. FID is reported as well as the improved Frechet distance with DINOv2 features (FDD), which appears to correlate with human judgements better than lnceptionV3 features. For video quality, FVD scores are also reported.
[0053] Symmetric epipolar distance (TSED) has been proposed as a more computationally efficient alternative to training separate NeRF models on every sample to evaluate 3D consistency. First, SIFT is used to obtain keypoints between a pair of views. For each keypoint in the first image, the epipolar line in the second image and its minimum distance to the corresponding keypoint can be computed. This is repeated for each keypoint in the second image to obtain a SED score. A hyperparameter threshold is selected, and the percentage of image pairs whose median SED is below the threshold is reported. TSED with a threshold of 2.0 is reported, and the TSED score of ground truth data is included, as poses in the data are noisy and do not achieve a perfect score.
[0054] A set of metrics named SfM distances is proposed. Here, COLMAP pose estimation is run on the generated views, and the camera extrinsics predicted by COLMAP are compared to the target poses. To get best results from COLMAP, the input views are also provided to it. The relative error in camera positions (can ) and the angular deviation of camera rotations in radians (can ) are reported. As with TSED scores, the SfM distances for the ground truth data are also reported. Due to the inherent scale and rotational ambiguity of COLMAP, additional care is taken to align the estimated poses to the original ones before comparing their differences as described above.
[0055] The aforementioned metrics for 3D consistency and pose alignment are invariant to the scale of the camera poses as they rely on epipolar distances and SfM. To evaluate the metric scale alignment, PSNR and SSIM are reported. These reconstruction-based metrics are not generally suitable to measure sample quality, as generative models can produce different butplausible samples. But for the same reason (i.e., that reconstruction metrics favor content alignment), they are instead found useful as indirect indicators of metric scale alignment. That is, given metric-scale poses, models are expected to not over / undershoot in position and rotation.
[0056] One common failure mode observed early is a tendency for certain models to copy input frames instead of generating temporal dynamics, and prior work has noted that good FVD scores can still be obtained with static video. A new metric called keypoint distance (KD) is therefore proposed, where the average motion of SIFT keypoints across image pairs with a sufficient number of matches is computed. Results are reported on generated and reference views to assess whether generated images have a similar motion distribution.
[0057] To calibrate RealEstatelOK, a single length scale per scene is predicted with the help of a zero-shot model for monocular depth that predicts metric depth from RGB and FOVTo calibrate RealEstatelOK, a single length scale per scene is predicted with the help of a zero-shot model for monocular depth that predicts metric depth from RGB and FOV. For each image, a metric depth is computed using the FOV provided by COLMAP. The length-scale is then obtained by regressing the 3D points obtained by COLMAP to metric depth at image locations to which the COLMAP points project. A loss based on the mean absolute error provides robustness to outliers in the depth map estimated by the metric depth models. The mean and variance of the per-frame scales yields a per-sequence estimator. The variance is used as a simple measure of confidence to identify scenes for which the scale estimate may not be reliable, discarding approximately 30% of scenes with the highest variance.
[0058] Both in-distribution and OOD evaluations for 3D NVS capabilities are considered in ablations and comparisons to prior work. The RealEstatelOK datasetis used as a common indistribution evaluation. To maximize the amount of training data, a specified percentage of the dataset is used as the validation split. Inference is run on all baselines for this split, noting that they might instead be advantaged as the test data may exist in their training data (for PNVS, all the test data is in fact part of their training dataset, giving them a distinct advantage). For OOD evaluation, the LLFF dataset can be used.
[0059] Main results on 3D novel view synthesis conditioned on a single image are presented herein, comparing capabilities against PhotoConsistentNVS2. These two are described as the strongest 3D scene NVS diffusion models with code and checkpoints available. For LLFF, the input trajectory of views is loaded in order, frames are subsampled evenly, and then a number of views are generated given a single input image. RealEstatelOK trajectories are much longer, so following PNVS, subsampling is performed with a specified stride. The NVS task using a specific number of views conditioned on a single image is used, in order to avoid weakening the baselines: MotionCtrl can only predict up to a certain number of frames, and PNVS performance degrades with the length of the sequence as it is an image-to-image model. Versions of the model that process a specific number of frames were trained for this end (as opposed to using the main 32-frame model and subsampling the output) for a more direct comparison. Quantitative metrics are computed on a number of scenes from the RealEstatelOK test split, and on all scenes in LLFF.
[0060] 4DiM can use a UViT architecture. Compared to the commonly employed UNet architecture, UViT uses a transformer backbone at the bottleneck resolution with no convolutions, leading to improved accelerator utilization. Information across frames mixes exclusively in temporal attention blocks. The temporal attention blocks can be employed to keep computational costs within reasonable limits compared to, for instance, 3D attention. Because the temporal attention blocks have a limited sequence length (e.g., 32 frames), they are inserted at all UViT resolutions. To limit memory usage, per-frame self-attention blocks can be used only at the bottleneck resolution (e.g., 16x16) where the sequence length is the product of the current resolution's dimensions.
[0061] The transformer blocks can combine various components. Parallel attention and MLP blocks can be used, as early experiments found that this leads to slightly better hardware utilization while achieving almost identical sample quality. Conditioning information can be injected, though fewer conditioning blocks can be used due to the use of the parallel attention + MLP. Query-key normalization can be additionally employed for improved training stability, but with RMSNorm for simplicity.
[0062] The same positional encodings for diffusion noise levels and relative timestamps can be used. For relative poses, following 3DiM, generation is conditioned with per-pixel ray origins and directions, as originally proposed by SRT in a regressive setting. This can be compared to conditioning via encoded extrinsics and focal lengths in a manner similar to Zero-l-to-3. While a thorough study of different encodings is missing in the literature, rays are a natural choice as they encode camera intrinsics in a way that is independent of the target resolution, yet gives the network precise, pixel-level information about what contents of the scene are visible and which are outside the field of view and therefore require extrapolation by the model. Unlike Plucker coordinates, rays additionally preserve camera positions which can be beneficial to deal with occlusions in the underlying 3D scene.
[0063] 4DiM models can be trained for a large quantity (e.g., IM) of steps. Using 64 TPU v5e chips, a throughput of approximately 1 step per second can be achieved with a batch size of 128. The model can include billions (e.g., 2.6B parameters), and available HBM can be maximized with various strategies such as: first, bfloatl6 activations are used (but still float32 weights to avoid instabilities); FSDP (i.e., zero-redundancy sharding with delayed and rematerialized all-gathers), and / or other strategy or strategies. This model can be then finetuned to its final 32-frame version. This strategy allows leveraging large batch-size pretraining, which is not possible with too many frames because the amount of activations stored for backpropagation scales linearly with respect to the number of frames per training example. The model can be finetuned with the same number of chips, albeit for a lesser quantity (e.g., 50,000) of steps and at a smaller batch size (e.g., 32). This allows sharding the frames of the video and using more than one chip per example. When using temporal attention, the keys and values are all-gathered and the queries are kept sharded over frames, though the computation of attention is not decomposed further as sufficient HBM is available to fully parallelize over all-gathered keys and values. FSDP is also enabled on the frame axis to maximize HBM savings so the number of shards is truly the number of chips (as opposed to the number of batchparallel towers). This is possible because frame sharding requires an all-reduce by mean on the loss over the frame axis. Identically to zero-redundancy sharding, this can be broken down into a reduce-scatter followed by an all-gather.
[0064] TSED can be computed only for contiguous pairs in each view trajectory, and sequences where less than 10 two-view keypoint matches are found are discarded. A threshold of 2.0 can be used in all experiments. For the proposed SfM distances, in order to obtain a scale-invariant metric, the camera positions predicted by COLMAP can be first aligned by relativizing the predicted and original poses with respect to the first conditioning frame. This should resolve rotation ambiguity. Then, scale ambiguity is resolved by analytically solving for the least squares optimization problem that aligns the original positions to re-scaled COLMAP camera positions. Finally, the relative error in positions under the L2 norm can be reported, i.e., normalization by the norm of the original camera positions is performed.
[0065] Tu rning now to the figures, Figure 1 illustrates an example image generation network architecture 100 that can be utilized in various implementations disclosed herein.
[0066] Input data 102A-N that can be processed utilizing the architecture 100 is illustrated and can include one or more images. The image(s) can be captured at corresponding first timestamp(s) and / or at corresponding first pose(s). The image(s) can include real image(s) captured by a real camera. The input data further includes one or more additional camera poses and / or one or more additional timestamp(s), that differ from the corresponding first timestamp(s). The architecture 100 can be used to process the input data and generate output data that includes one or more predicted images that are represented by 103A-N. The predicted image(s) each reflect scene(s) of the input image(s) but are at the additional camera pose(s) and / or the additional timestamp(s) reflected by the input data.
[0067] The image(s) of the input data can be processed using layers 140 and 141 to downsample the images to a lower resolution. The lower resolution image(s), or an encoding thereof, can be processed using a frame diffusion model with a frame (DiT) 150 and / or a time DiT 160. The resulting output can be image(s) that can be processed using layers 142 and 143 to up-sample the image(s) to higher resolution predicted image(s) 103A-N.
[0068] Figure 1 illustrates the frame DiT 150 in more detail, and the time DiT 160 can include similar architecture or can be integrated as part of the frame DiT 150. The frame DiT 150 includes FiLM / Identify layers 151A, 151B, and 151C which can function either as Feature-wise Linear Modulation (FiLM) layers or as identity functions, based on whether a conditioningsignal is provided. For example, "h" in Figure 1 illustrates conditioning signal(s) and the output from the FiLM layers 151A, 151B, and 151C can each correspond to a corresponding conditioning signal. For example, output eocan correspond to diffusion noise levels, output etcan correpond to timestamp(s), and ercan correspond to pose / per-pixel ray origins and directions.
[0069] The output from the FiLM layers 151A, 151B, and 151C can be injected as conditioning information into conditioning layers 153A and 153B. The projection layers 154A and 154B can perform dimensionality reduction or projection operations. The attention layer 155 can include self-attention and multi-layer perceptron (MLP) layer 156 can perform multi-layer perceptron operations. The norm layers 152A and 152B provide normalization.
[0070] Figure 2 shows a flowchart that depicts a method 200 for training an image generation network. The method can be implemented by one or more processors.
[0071] At block 202, the system identifies a plurality of training instances. Each training instance may include corresponding output of one or more corresponding ground truth images. Each training instance may also include corresponding training input that includes one or more corresponding conditioning images, a corresponding conditioning camera pose, and a corresponding conditioning timestamp. In some implementations, the training instances identified at block 202 can include pose-less instances that include real values for the corresponding conditioning images and for the corresponding conditioning timestamps, but include null or default values for the corresponding conditioning camera poses. In some implementations, the training instances identified at block 202 can include timestamp-less instances that include real values for the corresponding conditioning images and for the corresponding camera poses, but include null or default values for the corresponding conditioning timestamps. In some implementations, the training instances identified at block 202 can include full instances that may include real values for the corresponding conditioning images, for the corresponding camera poses, and for the corresponding conditioning timestamps.
[0072] At block 204, the system performs iterations of training the image generation network based on the training instances. Performing the iterations of training can include performingno-drop iterations in which the full training instances are utilized, performing camera pose drop iterations in which the pose-less training instances are utilized, and performing timestamp drop iterations in which the timestamp-less training instances are utilized.
[0073] In some implementations, block 204 can include preventing the pose-less instances from being utilized in the no-drop iterations and the timestamp drop iterations. In some implementations, block 204 can include preventing the timestamp-less instances from being utilized in the no-drop iterations and the camera pose drop iterations. In some implementations, block 204 can include causing the full instances to be utilized in the no-drop iterations. In some implementations, block 204 can include generating, in each of the iterations, a corresponding loss based on comparing the corresponding ground truth output to one or more predicted outputs generated using an image generation network (e.g. the architecture 100 of Figure 1). In some implementations, block 204 can include updating the image generation network based on the generated losses.
[0074] At block 206, the system, subsequent to performing the iterations of training, provides the trained image generation network for use by one or more computing devices. For example, the system can provide the trained image generation network for use in generating predicted images that reflect corresponding scenes of corresponding real images, but that have differing camera poses and / or are from differing points in time. In some of those implementations, such generated predicted images can be used in training a robotic policy and / or in control of a robot.
[0075] Figure 3 shows a flowchart that depicts a method 300. The method can be implemented by one or more processors.
[0076] At block 302, the system receives a real image of a scene. The real image can be captured at a first time by a real camera from a first camera pose of the real camera.
[0077] At block 304, the system receives a second camera pose that differs from the first camera pose and a timestamp that reflects a second time that is after the first time or that is before the first time.
[0078] At block 306, the system processes the real image, the second camera pose, and the timestamp, using an image generation network, to generate a predicted image that reflects the scene of the real image but that is from the second camera pose and that is at the second time.
[0079] FIG. 4 schematically depicts an example architecture of a robot 420. The robot 420 includes a robot control system 460, one or more operational components 440a-440n, and one or more sensors 442a-442m. The sensors 442a-442m may include, for example, vision sensors, light sensors, pressure sensors, pressure wave sensors (e.g., microphones), proximity sensors, accelerometers, gyroscopes, thermometers, barometers, and so forth. While sensors 442a-m are depicted as being integral with robot 420, this is not meant to be limiting. In some implementations, sensors 442a-m may be located external to robot 420, e.g., as standalone units.
[0080] Operational components 440a-440n may include, for example, one or more end effectors and / or one or more servo motors or other actuators to effectuate movement of one or more components of the robot. For example, the robot 420 may have multiple degrees of freedom and each of the actuators may control the actuation of the robot 420 within one or more of the degrees of freedom responsive to the control commands. As used herein, the term actuator encompasses a mechanical or electrical device that creates motion (e.g., a motor), in addition to any driver(s) that may be associated with the actuator and that translate received control commands into one or more signals for driving the actuator. Accordingly, providing a control command to an actuator may comprise providing the control command to a driver that translates the control command into appropriate signals for driving an electrical or mechanical device to create desired motion.
[0081] The robot control system 460 may be implemented in one or more processors, such as a CPU, GPU, and / or other controller(s) of the robot 420. In some implementations, the robot 420 may comprise a "brain box" that may include all or aspects of the control system 460. For example, the brain box may provide real time bursts of data to the operational components 440a-n, with each of the real time bursts comprising a set of one or more control commands that dictate, inter alia, the parameters of motion (if any) for each of one or more of theoperational components 440a-n. In some implementations, the robot control system 460 may perform one or more aspects of method(s) described herein, such as method 300 of FIG. 3.
[0082] As described herein, in some implementations all or aspects of the control commands generated by control system 460, in controlling a robot during performance of a robotic task, can be generated based on considering different poses and / or different time stamp image(s) generated according to techniques described herein. Although control system 460 is illustrated in FIG. 4 as an integral part of the robot 420, in some implementations, all or aspects of the control system 460 may be implemented in a component that is separate from, but in communication with, robot 420. For example, all or aspects of control system 460 may be implemented on one or more computing devices that are in wired and / or wireless communication with the robot 420, such as computing device 510.
[0083] FIG. 5 is a block diagram of an example computer system 510. Computer system 510 typically includes at least one processor 514 which communicates with a number of peripheral devices via bus subsystem 512. These peripheral devices may include a storage subsystem 524, including, for example, a memory subsystem 525 and a file storage subsystem 526, user interface output devices 520, user interface input devices 522, and a network interface subsystem 516. The input and output devices allow user interaction with computer system 510. Network interface subsystem 516 provides an interface to outside networks and is coupled to corresponding interface devices in other computer systems.
[0084] User interface input devices 522 may include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated into the display, audio input devices such as voice recognition systems, microphones, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information into computer system 510 or onto a communication network.
[0085] User interface output devices 520 may include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices. The display subsystem may include a cathode ray tube (CRT), a flat-panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The displaysubsystem may also provide non-visual display such as via audio output devices. In general, use of the term "output device" is intended to include all possible types of devices and ways to output information from computer system 510 to the user or to another machine or computer system.
[0086] Storage subsystem 524 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 524 may include the logic to perform selected aspects of method 200, method 300, and / or to implement one or more aspects of robot 400. Memory 525 used in the storage subsystem 524 can include a number of memories including a main random-access memory (RAM) 530 for storage of instructions and data during program execution and a read only memory (ROM) 532 in which fixed instructions are stored. A file storage subsystem 526 can provide persistent storage for program and data files, and may include a hard disk drive, a CD- ROM drive, an optical drive, or removable media cartridges. Modules implementing the functionality of certain implementations may be stored by file storage subsystem 526 in the storage subsystem 524, or in other machines accessible by the processor(s) 514.
[0087] Bus subsystem 512 provides a mechanism for letting the various components and subsystems of computer system 510 communicate with each other as intended. Although bus subsystem 512 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0088] Computer system 510 can be of varying types including a workstation, server, computing cluster, blade server, server farm, smart phone, smart watch, smart glasses, set top box, tablet computer, laptop, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computer system 510 depicted in FIG. 5 is intended only as a specific example for purposes of illustrating some implementations. Many other configurations of computer system 510 are possible having more or fewer components than the computer system depicted in FIG. 5.
[0089] Figures 6A, 6B, 6C, and 6D each present quantitative data organized in tables.
[0090] Figure 6A includes Table 1, which provides a comparison of 8-frame 4DiM models against prior work in 3D Novel Volume Synthesis (NVS). This table shows the performance ofmodels trained only on calibrated RealEstatelOK (4DiM-R) or with a full dataset mixture (4DiM), evaluated both in-distribution on the calibrated RealEstatelOK test set and in zero-shot conditions on the LLFF dataset. The metrics include FID, FDD, FVD, TSED, SfMD (position and rotation), LPIPS, PSNR, and SSIM.
[0091] Figure 6B includes Table 2, which shows an ablation study demonstrating the advantage of using calibrated data. This table compares 8-frame 4DiM models trained on calibrated RealEstatelOK versus uncalibrated RealEstatelOK and evaluates their performance on the RealEstatelOK, LLFF, and ScanNet++ datasets using the same set of metrics.
[0092] Figure 6C includes Table 3, which presents another ablation study highlighting the advantage of co-training with video data. This table compares an 8-frame 4DiM model trained with video data to an identical model trained without video data, evaluated on calibrated RealEstatelOK, ScanNet++, and calibrated LLFF datasets using the same metrics.
[0093] Figure 6D includes two tables. Table 4 compares different numbers of input frames used for video extrapolation with 32-frame 4DiM models, evaluating metrics such as FID, FDD, FVD, and KD. Table 5 presents quantitative performance metrics for generating panoramas in both extrapolation and interpolation regimes. This table shows the performance when using varying numbers of conditioning frames for Street View panoramas and Matterport3D 360° panoramas, using metrics including FID, FDD, PSNR, SSIM, and LPIPS.
[0094] In some implementations a method implemented by processor(s) is provided for generating a predicted image from a real image. The method involves receiving a real image captured at a first time from a first camera pose, and receiving a second camera pose that differs from the first camera pose and a timestamp that reflects a second time that is either after or before the first time. The real image, second camera pose, and timestamp are processed using an image generation network to generate a predicted image that reflects the scene of the real image from the second camera pose and at the second time.
[0095] In some implementations a method implemented by processor(s) is provided for training an image generation network (for processing an input image to generate an output image) to generate images based on conditioning images, camera poses, and timestamps. The method includes identifying training instances with supervised output of corresponding groundtruth images and corresponding training input that includes conditioning images, camera poses, and timestamps. Iterations of training are performed based on the training instances, including no-drop iterations, camera pose drop iterations, and timestamp drop iterations.
[0096] In some implementations, a method is provided that is a method of training an image generation network for use in generating, based on at least one conditioning image and a conditioning camera pose and / or a conditioning timestamp, a predicted image that reflects a scene of the conditioning image but that is from a differing camera pose reflected by the conditioning camera pose and / or that is from a differing point in time reflected by the conditioning timestamp. The method is implemented by processor(s) (e.g., graphics processing unit(s), tensor processing unit(s)) and includes identifying a plurality of training instances. Each of the training instances includes corresponding output of one or more corresponding ground truth images and corresponding training input that includes one or more corresponding conditioning images, a corresponding conditioning camera pose, and a corresponding conditioning timestamp. The method further includes performing iterations of training the image generation network based on the training instances. Performing the iterations of training includes: performing no-drop iterations in which nothing is dropped from the corresponding training inputs for the training instances used in the no-drop iterations; performing camera pose drop iterations in which only the corresponding conditioning camera pose, for the training instances used in the camera pose drop iterations, is dropped from the corresponding training inputs; and performing timestamp drop iterations in which only the corresponding conditioning timestamp, for the training instances used in the timestamp pose drop iterations, is dropped from the corresponding training inputs.
[0097] These and other implementations of the technology disclosed herein can include one or more of the following features.
[0098] In some implementations, the image generation network includes an image diffusion model and a plurality of chained feature-wise linear modulation (FiLM) layers. The chained FiLM layers including a conditioning image FiLM layer for processing conditioning images, a camera pose FiLM layer for processing conditioning camera poses, and a timestamp FiLM layer for processing conditioning timestamps. In some of those implementations: in the no-dropiterations none of the FiLM layers are configured to behave like the identity function; in the camera pose drop iterations only the camera pose FiLM layer is configured to behave like the identity function, thereby dropping the corresponding conditioning camera poses; and in the timestamp drop iterations only the timestamp FiLM layer is configured to behave like the identity function, thereby dropping the corresponding conditioning timestamps. In some versions of those implementations, configuring one of the FiLM layers to behave like the identity function includes setting a corresponding weight thereof to zero. In some of those or other versions, performing the iterations of training further includes performing camera, timestamp, and image drop iterations in which the corresponding conditioning images, the corresponding conditioning camera pose, and the corresponding timestamp, for the training instances used in the camera, timestamp, and image drop iterations, are dropped from the corresponding training inputs. Optionally, in the camera, timestamp, and image drop iterations the camera pose FiLM layer, the timestamp FiLM layer, and the conditioning image film layer are all configured to behave like the identity function, thereby dropping the corresponding conditioning images, the corresponding conditioning camera poses, and the corresponding timestamps.
[0099] In some implementations, the training instances include pose-less instances that include real values for the corresponding conditioning images and for the corresponding conditioning timestamps, but include null or default values for the corresponding conditioning camera poses. In some of those implementations, performing the iterations of training the image generation network based on the training instances includes preventing the pose-less instances from being utilized in the no-drop iterations and the timestamp drop iterations.
[0100] In some implementations, the training instances include timestamp-less instances that include real values for the corresponding conditioning images and for the corresponding camera poses, but include null or default values for the corresponding conditioning timestamps. In some of those implementations, performing the iterations of training the image generation network based on the training instances includes preventing the timestamp-less instances from being utilized in the no-drop iterations and the camera pose drop iterations.
[0101] In some implementations, the training instances include full instances that include real values for the corresponding conditioning images, for the corresponding camera poses, and for the corresponding conditioning timestamps. In some of those implementations, performing the iterations of training the image generation network based on the training instances includes causing the full instances to be utilized in the no-drop iterations.
[0102] In some implementations, performing the iterations of training includes generating, in each of the iterations, a corresponding loss based on comparing the corresponding ground truth output to one or more predicted outputs generated using the diffusion model, and updating the image generation network (e.g., at least the image diffusion model thereof) based on the loss. In some of those implementations, generating the corresponding loss includes determining an error between actual noise, that is based on the corresponding ground truth output, and predicted noise reflected in the one or more predicted outputs. Optionally, determining the error includes determining a mean absolute error, an LI loss.
[0103] In some implementations, the method further includes, subsequent to performing the iterations of training, providing the trained image generation network for use by one or more computing devices. In some versions of those implementations, providing the trained image generation network for use by one or more computing devices includes transmitting, via one or more communications network, the trained image generation network to the one or more computing devices. In those or other versions, providing the trained image generation network for use by one or more computing devices additionally or alternatively includes enabling access to the trained image generation network via an application programming interface.
[0104] In some implementations, the method further includes using the trained image generation network in generating predicted images that reflect corresponding scenes of corresponding real images, but that have differing camera poses and / or are from differing points in time. In some versions of those implementations, the method further includes using the predicted images in training a robotic policy. In some other versions of thoseimplementations, the corresponding real images are captured by a camera of a robot and the method further includes using the predicted images in control of the robot.
[0105] In some implementations a method implemented by processor(s) is provided and includes receiving an image of a scene. The image is captured at a first time at a first camera pose. The image can be a real image captured by a real camera at the first camera pose. The method further includes receiving a second camera pose that differs from the first camera pose and receiving a timestamp that reflects a second time that is after the first time or that is before the first time. The method further includes processing the real image, the second camera pose, and the timestamp, using an image generation network, to generate a predicted image. The generated predicted image reflects the scene of the real image but is from the second camera pose and is at the second time.
[0106] These and other implementations of the technology disclosed herein can include one or more of the following features.
[0107] In some implementations, the method further includes: receiving an additional image of an additional scene, where the additional image is captured at a third time from an additional camera pose; receiving a third camera pose that differs from the additional camera pose, but without receiving any additional timestamp that reflects a time that differs from the third time; and processing the additional image and the third camera pose, using the image generation network, to generate an additional predicted image that reflects the scene of the additional image but that is from the third camera pose. In some versions of those implementations, the image generation network includes an image diffusion model and a plurality of chained FiLM layers, the chained FiLM layers including a conditioning image FiLM layer for processing conditioning images, a camera pose FiLM layer for processing conditioning camera poses, and a timestamp FiLM layer for processing conditioning timestamps. In some of those implementations, in processing the additional image and the third camera pose, using the image generation network, the timestamp FiLM layer is configured to behave like the identity function and the timestamp FiLM layer is configured to behave like the identity function based on not receiving any additional timestamp that reflects a time that differs from the third time.
[0108] In some implementations, the method further includes: receiving a further image of a further scene, where the further image is captured at a fourth time at a further camera pose; receiving a further timestamp that reflects a fifth time that is after the fourth time or that is before the fourth time, but without receiving any additional camera pose that differs from the further camera pose; and processing the further real image and the further timestamp, using the image generation network, to generate a further predicted image that reflects the further scene of the further real image but that is at the fifth time.
[0109] In some versions of those implementations, the image generation network includes an image diffusion model and a plurality of chained FiLM layers, the chained FiLM layers including a conditioning image FiLM layer for processing conditioning images, a camera pose FiLM layer for processing conditioning camera poses, and a timestamp FiLM layer for processing conditioning timestamps. In some of those versions, in processing the further real image and the further timestamp, using the image generation network, the camera pose FiLM layer is configured to behave like the identity function and the camera pose FiLM layer is configured to behave like the identity function based on not receiving any additional camera pose that differs from the further camera pose.
[0110] In some implementations, the image generation network includes an image diffusion model and a plurality of chained FiLM layers, the chained FiLM layers including a conditioning image FiLM layer for processing conditioning images, a camera pose FiLM layer for processing conditioning camera poses, and a timestamp FiLM layer for processing conditioning timestamps. In some of those implementations, processing the real image, the second camera pose, and the timestamp, using the image generation network, includes processing the real image using the conditioning image FiLM layer, processing the second camera pose using the camera pose FiLM layer, and processing the timestamp using the timestamp FiLM layer.
[0111] In some implementations, the second camera pose and the timestamp are generated based on one or more user interface inputs at a client device.
[0112] In some implementations, the real camera is part of a robot, the second camera pose and the timestamp are generated by the robot, and the predicted image is used by the robot in determining whether to perform one or more robotic actions.
[0113] Some implementations includes a system having one or more processors and memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to perform one or more of the methods disclosed herein.
[0114] Some implementations includes a system having memory storing instructions and one or more processors operable to execute the instructions to cause performance of one or more of the methods disclosed herein.
[0115] Some implementations include at least one transitory or non-transitory computer-readable medium including instructions that, in response to execution by one or more processors, cause the one or more processors to perform one or more of the methods disclosed herein.
Claims
Claims1. A method of training an image generation network for use in generating, based on at least one conditioning image and a conditioning camera pose and / or a conditioning timestamp, a predicted image that reflects a scene of the conditioning image but that is from a differing camera pose reflected by the conditioning camera pose and / or that is from a differing point in time reflected by the conditioning timestamp, the method implemented by one or more processors and the method comprising: identifying a plurality of training instances each including: corresponding output of one or more corresponding ground truth images; and corresponding training input that includes one or more corresponding conditioning images, a corresponding conditioning camera pose, and a corresponding conditioning timestamp; performing iterations of training the image generation network based on the training instances, performing the iterations of training comprising: performing no-drop iterations in which nothing is dropped from the corresponding training inputs for the training instances used in the no-drop iterations; performing camera pose drop iterations in which only the corresponding conditioning camera pose, for the training instances used in the camera pose drop iterations, is dropped from the corresponding training inputs, and performing timestamp drop iterations in which only the corresponding conditioning timestamp, for the training instances used in the timestamp pose drop iterations, is dropped from the corresponding training inputs.
2. The method of claim 1, wherein the image generation network comprises an image diffusion model and a plurality of chained feature-wise linear modulation (FiLM) layers, the chained FiLM layers including a conditioning image FiLM layer for processing conditioning images, a camera pose FiLM layer for processing conditioning camera poses, and a timestamp FiLM layer for processing conditioning timestamps.
3. The method of claim 2,wherein in the no-drop iterations none of the FiLM layers are configured to behave like the identity function; wherein in the camera pose drop iterations only the camera pose FiLM layer is configured to behave like the identity function, thereby dropping the corresponding conditioning camera poses; and wherein in the timestamp drop iterations only the timestamp FiLM layer is configured to behave like the identity function, thereby dropping the corresponding conditioning timestamps.
4. The method of claim 3, wherein configuring one of the FiLM layers to behave like the identity function comprises setting a corresponding weight thereof to zero.
5. The method of claim 3 or claim 4, wherein performing the iterations of training further comprises performing camera, timestamp, and image drop iterations in which the corresponding conditioning images, the corresponding conditioning camera pose, and the corresponding timestamp, for the training instances used in the camera, timestamp, and image drop iterations, are dropped from the corresponding training inputs; and wherein in the camera, timestamp, and image drop iterations the camera pose FiLM layer, the timestamp FiLM layer, and the conditioning image film layer are all configured to behave like the identity function, thereby dropping the corresponding conditioning images, the corresponding conditioning camera poses, and the corresponding timestamps.
6. The method of any preceding claim, wherein the training instances include pose-less instances that include real values for the corresponding conditioning images and for the corresponding conditioning timestamps, but include null or default values for the corresponding conditioning camera poses.
7. The method of claim 6, wherein performing the iterations of training the image generation network based on the training instances comprises: preventing the pose-less instances from being utilized in the no-drop iterations and the timestamp drop iterations.
8. The method of any preceding claim, wherein the training instances include timestampless instances that include real values for the corresponding conditioning images and for the corresponding camera poses, but include null or default values for the corresponding conditioning timestamps.
9. The method of claim 8, wherein performing the iterations of training the image generation network based on the training instances comprises: preventing the timestamp-less instances from being utilized in the no-drop iterations and the camera pose drop iterations.
10. The method of any preceding claim, wherein the training instances include full instances that include real values for the corresponding conditioning images, for the corresponding camera poses, and for the corresponding conditioning timestamps.
11. The method of claim 10, wherein performing the iterations of training the image generation network based on the training instances comprises: causing the full instances to be utilized in the no-drop iterations.
12. The method of any preceding claim, wherein performing the iterations of training comprises: generating, in each of the iterations, a corresponding loss based on comparing the corresponding ground truth output to one or more predicted outputs generated using the diffusion model; and updating the image generation network based on the loss.
13. The method of claim 12, wherein generating the corresponding loss comprises determining an error between actual noise, that is based on the corresponding ground truth output, and predicted noise reflected in the one or more predicted outputs.
14. The method of claim 13, wherein determining the error comprises determining a mean absolute error.
15. The method of any preceding claim, further comprising, subsequent to performing the iterations of training: providing the trained image generation network for use by one or more computing devices.
16. The method of claim 15, wherein providing the trained image generation network for use by one or more computing devices comprises transmitting, via one or more communications network, the trained image generation network to the one or more computing devices.
16. The method of claim 15, wherein providing the trained image generation network for use by one or more computing devices comprises enabling access to the trained image generation network via an application programming interface.
17. The method of any preceding claim, further comprising: using the trained image generation network in generating predicted images that reflect corresponding scenes of corresponding real images, but that have differing camera poses and / or are from differing points in time.
18. The method of claim 17, further comprising: using the predicted images in training a robotic policy.
19. The method of claim 17, wherein the corresponding real images are captured by a camera of a robot and further comprising: using the predicted images in control of the robot.
20. A method implemented by one or more processors, the method comprising: receiving a real image of a scene, the real image being captured at a first time by a real camera from a first camera pose of the real camera; receiving a second camera pose that differs from the first camera pose and a timestamp that reflects a second time that is after the first time or that is before the first time; processing the real image, the second camera pose, and the timestamp, using an image generation network, to generate a predicted image that reflects the scene of the real image but that is from the second camera pose and that is at the second time.
21. The method of claim 20, further comprising: receiving an additional real image of an additional scene, the additional real image being captured at a third time by an additional real camera from an additional camera pose of the additional real camera;receiving a third camera pose that differs from the additional camera pose, but without receiving any additional timestamp that reflects a time that differs from the third time; processing the additional real image and the third camera pose, using the image generation network, to generate an additional predicted image that reflects the scene of the additional real image but that is from the third camera pose.
22. The method of claim 21, wherein the image generation network comprises an image diffusion model and a plurality of chained feature-wise linear modulation (FiLM) layers, the chained FiLM layers including a conditioning image FiLM layer for processing conditioning images, a camera pose FiLM layer for processing conditioning camera poses, and a timestamp FiLM layer for processing conditioning timestamps; and wherein in processing the additional real image and the third camera pose, using the image generation network, the timestamp FiLM layer is configured to behave like the identity function, wherein the timestamp FiLM layer is configured to behave like the identity function based on not receiving any additional timestamp that reflects a time that differs from the third time.
23. The method of any one of claims 20 to 22, further comprising: receiving a further real image of a further scene, the further real image being captured at a fourth time by a further real camera from a further camera pose of the further real camera; receiving a further timestamp that reflects a fifth time that is after the fourth time or that is before the fourth time, but without receiving any additional camera pose that differs from the further camera pose; processing the further real image and the further timestamp, using the image generation network, to generate a further predicted image that reflects the further scene of the further real image but that is at the fifth time.
24. The method of claim 23,wherein the image generation network comprises an image diffusion model and a plurality of chained feature-wise linear modulation (FiLM) layers, the chained FiLM layers including a conditioning image FiLM layer for processing conditioning images, a camera pose FiLM layer for processing conditioning camera poses, and a timestamp FiLM layer for processing conditioning timestamps; and wherein in processing the further real image and the further timestamp, using the image generation network, the camera pose FiLM layer is configured to behave like the identity function, wherein the camera pose FiLM layer is configured to behave like the identity function based on not receiving any additional camera pose that differs from the further camera pose.
25. The method of any one of claims 20 to 24, wherein the image generation network comprises an image diffusion model and a plurality of chained feature-wise linear modulation (FiLM) layers, the chained FiLM layers including a conditioning image FiLM layer for processing conditioning images, a camera pose FiLM layer for processing conditioning camera poses, and a timestamp FiLM layer for processing conditioning timestamps; and processing the real image, the second camera pose, and the timestamp, using the image generation network, comprises: processing the real image using the conditioning image FiLM layer; processing the second camera pose using the camera pose FiLM layer; processing the timestamp using the timestamp FiLM layer.
26. The method of any one of claims 20 to 25, wherein the second camera pose and the timestamp are generated based on one or more user interface inputs at a client device.
27. The method of any one of claims 20 to 25, wherein the real camera is part of a robot, wherein the second camera pose and the timestamp are generated by the robot, and wherein the predicted image is used by the robot in determining whether to perform one or more robotic actions.
28. A system comprising one or more processors and memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to perform the method of any one of claims 1-27.
29. At least one non-transitory computer-readable medium comprising instructions that, in response to execution by one or more processors, cause the one or more processors to perform the method of any one of claims 1-27.