Panoramic video generation

The method employs a generative machine learning model to synthesize missing regions in panoramic videos, addressing the challenges of camera movement and dynamic objects, resulting in high-resolution and seamless panoramic videos.

JP2026507312AActive Publication Date: 2026-03-02GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025539981
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-27
Publication Date
2026-03-02
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

Conventional systems fail to create panoramic videos due to missing portions caused by camera movement and scene motion, especially with dynamic objects, making it impossible to stitch panoramic videos effectively.

Method used

A computer-implemented method using a generative machine learning model to reproject input panning videos onto a panoramic canvas, generate synthesized pixels for missing regions, and refine the video through temporal upsampling and spatial alignment, merging with input pixels to create a seamless panoramic video.

Benefits of technology

Generates high-quality panoramic videos with improved resolution and accurate motion tracking of dynamic objects by synthesizing missing regions, enhancing the compositional integrity of panoramic videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026507312000001_ABST
    Figure 2026507312000001_ABST
Patent Text Reader

Abstract

The media application receives an input panning video including image frames captured by a camera that pans at least once from a first side to a second side. The media application further includes reprojecting the input panning video onto a panoramic canvas, the panoramic canvas including the image frames and the missing region. The media application further includes generating a base video including synthesized pixels of the missing region at a coarse time scale, and outputting the panoramic video including synthesized pixels of the missing region of the panoramic canvas using a generative machine learning model by performing temporal upsampling, merging the temporally upsampled base video with the panoramic canvas, and refining the base video by resynthesizing the synthesized pixels of the missing region in the base video.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is an international application claiming priority under 35 U.S.C. §119(e) to U.S. Provisional Patent Application No. 63 / 616,037, entitled "Generating Panoramic Videos," filed December 29, 2023, which is incorporated herein by reference in its entirety. [Background technology]

[0002] Panorama stitching refers to simulating an image of a wider field of view from a set of images captured by rotating a camera in place. For example, a set of images can be stitched together when the camera moves 180 degrees along the same axis. Panorama stitching includes sparse feature-based or semi-dense registration, rigid or depth-compensated alignment, and image fusion to resolve seams and inconsistencies across views.

[0003] Panoramic videos are more complex than panoramic images because scene motion can cause failures in both alignment and composition, especially when the panoramic video contains dynamic objects. Furthermore, it is impossible to create a panoramic video from images using conventional systems because when a camera captures an image from a first position along an axis, other images are not captured from other positions along the axis, resulting in portions of the panoramic video being missing.

[0004] The description of the background art provided herein is intended to provide a general context for the present disclosure. To the extent provided in this background art section, the work of the inventors named herein, and aspects of the present disclosure that may not qualify as prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure. Summary of the Invention

[0005] The computer-implemented method includes receiving an input panning video including image frames captured by a camera that pans at least once from a first side to a second side. The method further includes reprojecting the input panning video onto a panoramic canvas, the panoramic canvas including the image frames and the missing region. The method further includes generating a base video including synthesized pixels of the missing region at a coarse time scale, and outputting the panoramic video including synthesized pixels of the missing region of the panoramic canvas using a generative machine learning model by performing temporal upsampling, merging the temporally upsampled base video with the panoramic canvas, and refining the base video by resynthesizing the synthesized pixels of the missing region in the base video. In some embodiments, the computer-implemented method is executed as a media application on a computing device.

[0006] In some embodiments, temporal upsampling includes temporally upsampling the base video and performing spatial and color alignment of the temporally upsampled base video with the panoramic canvas to obtain an aligned, temporally upsampled base video. In some embodiments, merging the temporally upsampled base video with the panoramic canvas includes merging the aligned, temporally upsampled base video with the panoramic canvas. In some embodiments, recombining includes generating a mask of pixels of the missing region, regenerating synthesized pixels of the missing region based on the mask, and merging the regenerated pixels with the synthesized pixels. In some embodiments, the quality of the panoramic video improves as a function of the number of times image frames captured by the camera include panning from a first side to a second side and panning from the second side to the first side. In some embodiments, the input panning video includes one or more moving objects. In some embodiments, the method further includes, in response to outputting the panoramic video, applying a super-resolution pass to the panoramic video to increase the resolution of the panoramic video. In some embodiments, the method further includes selecting a panoramic image from the panoramic video.

[0007] In some embodiments, a non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more computers, cause the one or more computers to perform operations, including receiving an input panning video including image frames captured by a camera panning at least once from a first side to a second side, and reprojecting the input panning video onto a panoramic canvas, the panoramic canvas including the image frames and a missing region, the operations further including generating a base video at a coarse time scale including pixels of the missing region, and outputting the panoramic video including the synthesized pixels of the missing region of the panoramic canvas using a generative machine learning model by performing temporal upsampling, merging the temporally upsampled base video with the panoramic canvas, and refining the base video by resynthesizing the synthesized pixels of the missing region in the base video.

[0008] In some embodiments, the temporal upsampling includes temporally upsampling the base video and performing spatial and color alignment of the temporally upsampled base video with the panoramic canvas to obtain an aligned temporally upsampled base video. In some embodiments, merging the temporally upsampled base video and the panoramic canvas includes merging the aligned temporally upsampled base video with the panoramic canvas. In some embodiments, the recombining includes generating a mask of pixels of the missing region, regenerating synthesized pixels of the missing region based on the mask, and merging the regenerated pixels with the synthesized pixels. In some embodiments, the input panning video includes one or more moving objects. In some embodiments, the operations further include, in response to outputting the panoramic video, applying a super-resolution pass to the panoramic video to increase the resolution of the panoramic video. In some embodiments, the operations further include selecting a panoramic image from the panoramic video.

[0009] In some embodiments, a system includes one or more processors and a memory coupled to the one or more processors, the memory storing instructions that, when executed by the processor, cause the processor to perform operations, including receiving an input panning video including image frames captured by a camera panning at least once from a first side to a second side, and reprojecting the input panning video onto a panoramic canvas, the panoramic canvas including the image frames and the missing region, the operations further including generating a base video at a coarse time scale including pixels of the missing region, and outputting the panoramic video including the synthesized pixels of the missing region of the panoramic canvas using a generative machine learning model by performing temporal upsampling, merging the temporally upsampled base video with the panoramic canvas, and refining the base video by resynthesizing the synthesized pixels of the missing region in the base video.

[0010] In some embodiments, the temporal upsampling includes temporally upsampling the base video and performing spatial and color alignment of the temporally upsampled base video with the panoramic canvas to obtain an aligned temporally upsampled base video. In some embodiments, merging the temporally upsampled base video and the panoramic canvas includes merging the aligned temporally upsampled base video with the panoramic canvas. In some embodiments, the recombining includes generating a mask of pixels of the missing region, regenerating synthesized pixels of the missing region based on the mask, and merging the regenerated pixels with the synthesized pixels. In some embodiments, the input panning video includes one or more moving objects. In some embodiments, the operations further include, in response to outputting the panoramic video, applying a super-resolution pass to the panoramic video to increase the resolution of the panoramic video. In some embodiments, the operations further include selecting a panoramic image from the panoramic video. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram of an exemplary network environment, according to some embodiments described herein. [Figure 2] FIG. 1 is a block diagram of an exemplary computing device according to some embodiments described herein. [Figure 3] 1 illustrates input images from an input panning video, the input images projected onto a panoramic canvas, and the resulting output images from an output panoramic video, according to some embodiments described herein. [Figure 4] 1 illustrates an example token-based model for outputting panoramic video, according to some embodiments described herein. [Figure 5]1 illustrates an example diffusion model trained to output a panoramic video, according to some embodiments described herein. [Figure 6A] 1 illustrates an example process for generating a panoramic video according to some embodiments described herein. [Figure 6B] 1 illustrates an example process for generating a panoramic video according to some embodiments described herein. [Figure 6C] 1 illustrates an example process for generating a panoramic video according to some embodiments described herein. [Figure 6D] 1 illustrates an example process for generating a panoramic video according to some embodiments described herein. [Figure 7] 1 illustrates an example of how a diffusion model performs upsampling and outpainting to generate panoramic video, according to some embodiments described herein. [Figure 8] 10 shows an exemplary comparison of how probabilities during outpainting use a token-based model compared to a diffusion model, according to some embodiments described herein. [Figure 9] 1 illustrates a flowchart of an example method for generating panoramic video, according to some embodiments described herein. DETAILED DESCRIPTION OF THE INVENTION

[0012] overview When a user captures input video by panning a camera left and right (or in any direction to capture different areas), the resulting panoramic video may have more composited pixels than the original pixels. This further increases the need for image generation that properly tracks the movement of objects in the panoramic video. For example, a panoramic video of people kayaking on a river requires not only image generation of stationary objects, such as buildings along the horizon, but also tracking and aligning movements, such as the movement of a paddle in the water. The task is also more difficult when the camera capturing the input video is moving during capture.

[0013] The techniques described herein advantageously generate panoramic videos by reprojecting an input panning video onto a panoramic canvas, where the panoramic canvas includes image frames from the input panning video and missing regions. A generative machine learning model outputs a panoramic video including synthesized pixels of the missing regions of the panoramic canvas. For example, the generative machine learning model may generate a base video at a coarse time scale including synthesized pixels of the missing regions. In some embodiments, the generative model (and / or other techniques) may refine the base video by performing temporal upsampling, merging the temporally upsampled base video with input pixels from the panoramic canvas, and recomposing synthesized pixels of the missing regions. The techniques may be used with videos including moving objects, moving people, portrait panoramic pans, etc.

[0014] Exemplary Computing Device 2 is a block diagram of an example computing device 200 that may be used to implement one or more features described herein. Computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, computing device 200 is media server 101. In another example, computing device 200 is user device 115.

[0015] In some embodiments, computing device 200 includes a processor 235, a memory 237, an input / output (I / O) interface 239, a display 241, a camera 243, and a storage device 245, all coupled via a bus 218. Processor 235 may be coupled to bus 218 via signal line 222, memory 237 may be coupled to bus 218 via signal line 224, I / O interface 239 may be coupled to bus 218 via signal line 226, display 241 may be coupled to bus 218 via signal line 228, camera 243 may be coupled to bus 218 via signal line 230, and storage device 245 may be coupled to bus 218 via signal line 232.

[0016] Processor 235 may be one or more processors and / or processing circuits that execute program code and control basic operations of computing device 200. A “processor” includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), a system having multiple processing units (e.g., in a multiprocessor configuration), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), dedicated circuitry for implementing functionality, a dedicated processor for performing processing based on neural network models, neural circuitry, a processor optimized for matrix calculations (e.g., matrix multiplication), or other systems. In some embodiments, processor 235 may include one or more coprocessors that perform neural network processing. In some embodiments, processor 235 may be a processor that processes data to produce a probabilistic output; for example, the output produced by processor 235 may be inaccurate or accurate within a range from an expected output. Processing need not be limited to a particular geographic location or have time limitations. For example, a processor may perform its functions in real time, offline, in batch mode, etc. Portions of processing may be performed at different locations at different times by different (or the same) processing systems. A computer may be any processor in communication with a memory.

[0017] Memory 237 is typically provided in computing device 200 for access by processor 235 and may be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc., located separately from and / or integrated with processor 235, suitable for storing instructions for execution by the processor or set of processors. Memory 237 may store software operated on computing device 200 by processor 235, including media application 103.

[0018] Memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, an image library application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, etc. One or more methods disclosed herein may operate in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having web pages, as a mobile application (“app”) running on a mobile computing device, etc.

[0019] Application data 266 may be data generated by other applications 264 or by hardware of computing device 200. For example, application data 266 may include images used by an image library application, user actions identified by other applications 264 (e.g., social networking applications), etc.

[0020] I / O interface 239 may provide functionality that allows computing device 200 to interface with other systems and devices. The interfaced devices may be included as part of computing device 200 or may be separate and in communication with computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 245), and input / output devices may communicate through I / O interface 239. In some embodiments, I / O interface 239 may be connected to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone, scanner, sensor, etc.) and / or output devices (display device, speaker device, printer, monitor, etc.).

[0021] Some examples of interface devices that can be connected to I / O interface 239 can include display 241, which can be used to display content, e.g., images, video, and / or user interfaces of output applications described herein, and to receive touch (or gesture) input from a user. For example, display 241 can be utilized to display a user interface. Display 241 can include any suitable display device, such as a liquid crystal display (LCD), light-emitting diode (LED), or plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touchscreen, a three-dimensional display screen, or other visual display device. For example, display 241 can be a flat display screen provided on a mobile device, multiple display screens embedded in an eyeglass form factor or headset device, or a monitor screen of a computing device.

[0022] Camera 243 may be any type of image capture device capable of capturing images and / or video. In some embodiments, camera 243 captures images or video that I / O interface 239 sends to media application 103.

[0023] The storage device 245 stores data related to the media application 103. For example, the storage device 245 may store a training dataset including training data such as a plurality of training images, a generative machine learning model, and an outpainter model.

[0024] The media application 103 receives input panning video, which includes image frames captured by a camera 243 panning at least once from a first side to a second side. In some embodiments, the camera 243 captures multiple pans in a single capture, such as panning from a first side to a second side and then back to the first side again. For example, a user may take a user device 115 and move it left and right multiple times to capture input video. The resolution of the panoramic video is increased by multiple pans to improve tracking of motion at different spatial locations.

[0025] The resolution of the panoramic video increases by the number of image frames captured at different angles, because as the number of image frames increases, the number of synthesized pixels generated for missing areas in the panoramic video decreases. In some embodiments, the input panning video includes one or more moving objects, such as people in motion (e.g., people kayaking in a river, people scuba diving, people walking around a landmark, people skating), cars, trees, water, etc.

[0026] The media application 103 reprojects the input panning video onto the panoramic canvas to map the input frame (x 0 ) and the corresponding mask frame of valid pixels (m 0) The mask can be used to identify valid pixels to be retained and missing areas where composited pixels will be generated. The panoramic canvas is a canvas with the dimensions of the panorama that contains the image frames and missing areas (of the scene not captured by the camera in the image frames captured at that moment). The panorama canvas contains both the image frames from the input panning video positioned based on their coordinate locations in the panoramic canvas, and the missing areas, which may be depicted as blank pixels, black pixels, or no data.

[0027] 3 shows three input images 300 from an input panning video, the three input images 300 projected onto a panoramic canvas 310, and three output images 320 from an output panoramic video. The input images 300 show people kayaking on a river. A first input image 300a includes a person kayaking 301 in the foreground, a person kayaking 302 in the background, and a horizon 303. A second input image 300b includes the first person 301 at a later point in the input panning video, with a paddle 304 in a different position compared to the first input image 300a, and a kayak 305 also visible. A third input image 300c includes a second person 302 in a different position compared to the first input image 300a.

[0028] The panoramic canvas 310 includes each input image 300 reprojected onto the panoramic canvas 310. In some embodiments, the media application 103 uses a projective transformation solver to perform the reprojection of the input panning video, which includes rotation but not translation or scaling. In response to determining that parallax correction is required, the media application 103 may use simultaneous localization and mapping (SLAM) to perform the reprojection. In some embodiments, the media application 103 ignores the translation of the cameras 243, calculates the altitude and azimuth of the ray for each input pixel relative to the direction of the corresponding camera 243, and then projects each ray onto the equirectangular canvas.

[0029] The output image 320 includes composited pixels of the missing areas of the panoramic canvas 310. For example, the second output image 320b includes a second person 302 that was not visible in the second input image 300b.

[0030]

number

[0031] In some embodiments, the input panning video is wider than the native aspect ratio of the generative machine learning model. The media application 103 scales the input panning video (x K ) can be downscaled.

[0032] Generative Machine Learning Models The generative machine learning model generates a base video that includes synthesized pixels of the missing regions at a coarse time scale. In some embodiments, the generative machine learning model outputs multiple overlapping spatial windows that span the panoramic width. This process is described in more detail below with reference to Figures 6A-6D, 7, and 8. The generative machine learning model may average the distributions that the generative machine learning model predicts over each window and output base image frames of the base video based on the averages.

[0033] The generative machine learning model refines the base video using temporal upsampling, merging the base video with input pixels from the input panning video, and recomposing the synthesized pixels of missing regions.

[0034] The media application 103 uses the training data to train a generative machine learning model to output panoramic video from input data, such as application data 266 (e.g., input panning video captured by the user device 115). The training data may come from any source, such as a data repository explicitly designated for training, data that has been given permission to be used as training data for machine learning, etc. In some embodiments, training may occur on the media server 101, which provides the training data directly to the user device 115, or training may occur locally on the user device 115, or a combination of both.

[0035] The trained machine learning model may include one or more model forms or structures, such as a linear network, a deep learning neural network implementing multiple layers (e.g., a "hidden layer" between an input layer and an output layer, where each layer is a linear network), a convolutional neural network (e.g., a network that divides or partitions input data into multiple portions or tiles, processes each tile separately using one or more neural network layers, and aggregates the results of processing each tile), a sequence-to-sequence neural network (e.g., a network that receives sequence data as input, such as words in a sentence or frames of a video, and produces a resulting sequence as output), or any other type of neural network.

[0036] The model format or structure may specify the connectivity between various nodes and their organization into layers. For example, nodes in a first layer (e.g., input layer) may receive data as input data or application data 266. Such data may include, for example, one or more pixels per node, for example, when the trained model is used, for example, to analyze images. Subsequent intermediate layers may receive as input the outputs of nodes in the previous layer according to the connectivity specified in the model format or structure. These layers may also be referred to as hidden layers. The final layer (e.g., output layer) produces the output of the generative machine learning model. In some implementations, the model format or structure also specifies the number and / or type of nodes in each layer.

[0037] In different embodiments, the trained generative machine learning model may include one or more models. One or more of the models may include multiple nodes arranged in layers according to a model structure or format. In some embodiments, a node may be, for example, a memoryless computational node configured to process one unit of input and generate one unit of output. The computation performed by the node may include, for example, multiplying each of multiple node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce the node output. In some embodiments, the computation performed by the node may also include applying a step function / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computation may include operations such as matrix multiplication. In some embodiments, computations by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, individual processing units of a graphics processing unit (GPU), or dedicated neural circuitry. In some implementations, a node may include memory, e.g., be capable of storing and using one or more previous inputs when processing a subsequent input. For example, a node with memory may include a long short-term memory (LSTM) node. An LSTM node may use memory to maintain a "state" that allows the node to function like a finite state machine (FSM).

[0038] In some implementations, the trained model may include embeddings or weights for individual nodes. For example, the model may begin as a number of nodes organized into layers, as specified by the model format or structure. At initialization, corresponding weights may be applied to nodes connected according to the model format, e.g., to the connection between each pair of nodes in successive layers of a neural network. For example, each weight may be randomly assigned or initialized to a default value. The model may then be trained, e.g., using training data, to produce results.

[0039] The training may include applying supervised learning techniques. In supervised learning, the training data may include multiple inputs (e.g., multiple input panning videos) and corresponding ground truth outputs for each input (e.g., a ground truth panoramic video for each input panning video). Based on a comparison of the model's output (e.g., the panoramic video) and the ground truth output (e.g., the ground truth panoramic video), the values ​​of the weights are automatically adjusted, for example, in a manner that increases the probability that the model generates a ground truth panoramic video.

[0040] In various implementations, the trained model includes a set of weights, or embeddings, corresponding to the model structure. In some implementations, the trained generative machine learning model may include an initial set of weights, for example, downloaded from a server that provides the weights. In various implementations, the trained generative machine learning model includes a set of weights, or embeddings, corresponding to the model structure. In implementations where data is omitted, the media application 103 may generate a trained generative machine learning model based on prior training, for example, by the developer of the media application 103, by a third party, etc.

[0041] In some embodiments, if the generative machine learning model includes a convolutional neural network trained using supervised learning, training the generative machine learning model may include, for each input panning video, obtaining a panoramic video based on the input panning video. The generative machine learning model may calculate a loss value based on a comparison between the panoramic video and a ground truth panoramic video (included in the training data). For example, the generative machine learning model may calculate an optical flow endpoint error (EPE) to measure the consistency of the generated motion. The flow between consecutive frames of the ground truth video and the output video may be calculated as an L2 difference. In some embodiments, small motion at low image resolutions may be evaluated based on a grid-based flow to generate more reliable sub-pixel alignment than a network-based flow. In some embodiments, pixel-level metrics may be evaluated separately for static and dynamic regions.

[0042] The generative machine learning model may update the weights of one or more nodes of the convolutional neural network based on the loss value (e.g., so that the loss value decreases until it falls below a threshold value after performing another cycle of tuning and training).

[0043] In some embodiments, the generative machine learning model is a token-based model. FIG. 4 shows an example token-based model 400 that outputs a panoramic video, according to some embodiments described herein. The token-based model 400 includes an encoder 410, a bidirectional transformer 420, and a decoder 430. The encoder 410 receives a downsampled masked video 405 (e.g., a reprojected input panning video) (e.g., to 160x96 pixels at 11 frames per second) and compresses the masked video 405 into masked tokens 415. In some embodiments, the mask includes input pixels where synthesized pixels are generated for areas outside the mask. The masked tokens 415 represent discrete embeddings that describe the content of the masked video 405. In some embodiments, the attention layer of the encoder 410 is masked so that later frames attend to earlier frames within a window, but earlier frames cannot attend to later frames.

[0044] The bidirectional transformer 420 iteratively transforms the masked tokens 415 into video tokens 425. ... K ) to reduce or remove artifacts in the panoramic video. The bidirectional transformer 420 may take the results of the forward pass within the region that the bidirectional transformer 420 previously examined and combine the results with the backward pass over the remainder of the region to obtain a video token 425. The video token 425 is decoded by the decoder 430 to output a similarly downsampled (e.g., 160x96 pixels at 11 frames per second) panoramic video 435. The panoramic video 435 is then upsampled (e.g., to 320x192 pixels).

[0045] In some embodiments, the generative machine learning model is a diffusion model, such as a spatiotemporal pixel diffusion model. The diffusion model is trained to generate a panoramic video by gradually adding noise to the panoramic video (noise addition), then performing a denoising process, and training the diffusion model to recover the original panoramic video from the noise. The media application 103 receives the reprojected input panning video as input and trains the diffusion model to output a panoramic video that includes synthesized pixels for the missing portions of the panoramic canvas.

[0046] In some embodiments, the media application 103 trains the diffusion model on an image outpainting task, where the training data includes image pairs: an image portion of a ground truth image and a corresponding image in which a random pixel or group of pixels has been removed. As a result of training the diffusion model on the outpainting task, the diffusion model is trained to receive an incomplete input image and output an image with synthesized pixels that replace any missing regions.

[0047] In some embodiments, the diffusion model is fine-tuned using training videos in which fixed masks are replaced with panning masks to simulate the situation where a user captures input panning video by moving the camera left and right. In some embodiments, the media application 103 performs fine-tuning of the trained diffusion model using video pairs of natural videos (ground truth videos) paired with masked videos with synthetic panoramic masks designed to mimic real masks.

[0048] 5 illustrates an exemplary diffusion model 500 trained to output a panoramic video, according to some embodiments described herein. The diffusion model 500 receives noise 505 and masked video 510. In some embodiments, the diffusion model 500 receives downsampled noise 505 and masked video 510 (e.g., 80 frames of 128x128 pixels) and outputs a similarly downsampled panoramic video 520. The panoramic video 520 may be a base layer that is upsampled to 1024x1024 pixels.

[0049] 6A-6D illustrate an exemplary process for generating a panoramic video. While these examples refer to a particular number of frames, the number of frames is used as an example and other numbers of frames may be used. FIG. 6A includes a process 600 using an input panning video 605. A panoramic projection 610 projects the input panning video 605 onto a panoramic canvas (the finest temporally detailed input frame x). 0 The spatial downsampling 615 generates a downsampled reprojected input video 620a (starting from the temporally coarse input frame x K , which ends with . The downsampled reprojected input video 620a may be progressively downsampled to match the native height of the generative machine learning model 625. For example, each level of the reprojected input video 620 may be associated with 2X downsampling. Although four levels are illustrated, any number of levels (k) may be used, such as two to five levels, or more than five levels. In some embodiments, the time scale k(x k ) is an input video at time scale k(m k ) has a corresponding mask.

[0050] Two different sliding space windows 621, 622 are used to provide portions of the reprojected base input video 620a to the generative machine learning model 625. The sliding space windows 621, 622 may be associated with dimensions that vary depending on the type of generative machine learning model 625 used to generate the base video 630a. The sliding space windows 621, 622 advantageously allow the process 600 to be used with input videos 605 of various panoramic widths.

[0051] The generative machine learning model 625 generates the base video 630a by performing spatial outpainting to generate synthesized pixels for missing regions. The model predictions within the sliding spatial window 631, 632 are averaged and new samples are drawn from the average.

[0052] In embodiments where the generative machine learning model 625 is a diffusion model, the projected panoramic canvas is cropped to each input window, and missing regions outside the image frame are outpainted. The Gaussian distributions (represented by μ and Σ) over pixel values ​​may be averaged, and new samples may be extracted using a diffusion method such as a denoising probabilistic model (DDPM).

[0053] In embodiments where the generative machine learning model 625 is a token-based model, spatial aggregation may be performed by averaging the predicted probability distributions for the tokens before sampling.

[0054] The generative machine learning model 625 gradually restores the temporal dimension of the input video 605 by gradually completing the panoramic video at an upsampled temporal resolution. For each level k∈[K−1,0], the generative machine learning model 625 calculates the reprojected input video 620 (x k ) to create a coarser panoramic video (y k+1 ) and the completed low-resolution panoramic video 630(y k )

[0055]

number

[0056] During this process, the diffusion sampling is constrained to maintain the temporally upsampled and synthesized pixels in odd-numbered frames, while the synthesized pixels are resynthesized in even-numbered frames. For fast-motion videos (e.g., skiers, kayakers, skateboarders, etc., or any scene where the subject's motion meets a threshold rate of change between consecutive frames of the video), the temporally upsampled pixels in odd-numbered frames may also need to be resynthesized. In that case, a full-frame mask is maintained in odd-numbered frames for a predetermined amount (e.g., 1 / 8) of the sampling schedule, and then the original mask (m) for the entire window is rescaled so that the diffusion model synthesizes new temporal details in the missing areas outside the original mask. k ) is used.

[0057] During outpainting 720, in the temporal dimension, a diffusion model 734 is applied in a sliding window manner with half of the window overlapping (e.g., 40 frames) between the two windows 730, 732. In the spatial dimension, multiple overlapping predictions are computed in parallel by the diffusion model 734 and then aggregated (736) to produce the completed panoramic video (y k )725 is extracted from the average.

[0058] In some embodiments, the diffusion model 734 is fine-tuned by randomly cropping the video to 128x128 pixels and augmenting it with a set of synthetic and actual panning video masks. In some embodiments, the diffusion model 734 is optimized for diffusion denoising objects with a squared error loss using dynamic masks (i.e., masks that change position in subsequent frames to simulate the input panning video).

[0059] Compared to the diffusion model, which performs outpainting, the token-based model performs masking (m k ) outpaint pixels outside.

[0060]

number

[0061] In some embodiments, the token-based model is fine-tuned to reduce fidelity loss in the encoder / decoder architecture to better align synthesized pixels with the input video and preserve input detail in outpainted regions. The decoder can be fine-tuned on patches of valid pixels before synthesizing the final result.

[0062] Referring to Figure 8, an exemplary comparison 800 of how probabilities during outpainting using a token-based model compare to a diffusion model is shown, according to some embodiments described herein. A reprojected input video 801 includes a missing region 802. A generative machine learning model generates samples 803 from two predictive probability distributions associated with two overlapping windows, namely, a left window 804 and a right window 805. The token-based model outputs discrete token probabilities 806, 807 based on the left window 804 and the right window 805, respectively, which are aggregated 808. The diffusion model outputs pixel probabilities 809, 810 based on Gaussian distributions for pixel values ​​(represented by μ and Σ), which are aggregated 811.

[0063] Once the final upsampled panoramic video 630d is generated, spatial super-resolution 635 is applied to the upsampled panoramic video 630d, and the result is merged with the output from reprojecting the input panning video 605 to obtain the full-resolution panoramic video 645.

[0064] 6B shows an example number of frames at each stage where the generative machine learning model is a token-based model 656 that accepts 11 frames 655. The input video 651 contains 35 frames. For the coarsest time scale, the reprojected input video is complete at 11 frames 655. 2X upsampling of 11 frames is 22 frames 654, 2X upsampling of 22 frames is 44 frames 653, and the 44 frame version is resampled to 35 frames 652.

[0065] 6C shows an example process 660 for converting an input video 661 into a panoramic video. The input video is a video of a person kayaking. A dashed box 662 in the input video highlights a section of the input video 661 where unobserved motion occurs during the capture of the panning video. For example, if one person's kayak moves from left to right and another person's kayak moves from right to left, the activity within the dashed box 662 will not be visible to the camera.

[0066] One difficulty in generating a panoramic video is ensuring that movements, such as a person using a paddle, are properly aligned. The reprojected input video 663 includes unknown regions 664 where portions of the reprojected input video 663 must be filled in with synthesized pixels to create the panoramic video. The movement of the person with the paddle in the reprojected input video 663 goes from rest to paddling, then back to rest again. If the synthesized pixels generated by the generative machine learning model 665 are not accurate, various issues (e.g., discontinuous, jerky, or unnatural movement of the paddle) may occur instead of the expected smooth paddle movement.

[0067] The generative machine learning model 665 outputs a panoramic video 666 in which missing regions are replaced with synthetic pixels and the synthetic image is merged with the reprojected input video 663. For example, original pixels in unmasked regions in the reprojected input video 663 are merged with the synthetic image.

[0068] FIG. 6D illustrates a process 670 for finalizing a panoramic video using a token-based model 679. In this example, a low-resolution panoramic video 671, which includes 11 frames with upsampled temporal context in a reprojected input video 672, is upsampled along with the reprojected input video 672, which includes 22 frames and temporal detail, resulting in a 2x upsampled video 673 with 22 frames. The upsampled video is then modified, for example, using the process described in FIG. 6A, by performing spatial alignment 674 (e.g., using temporal upsampling and frame rate matching), color alignment 675 (e.g., through color histogram matching), and compositing 676 (e.g., using grid-warp-based optical flow or other types of image interpolation and / or blending), resulting in a merged video 677 with 22 frames. In this example, a token-based model is used to generate the composited pixels. In some embodiments in which a different type of generative machine learning model is used, such as a diffusion model, color alignment may not be performed.

[0069] The token-based model 679 may receive as input a merged video 677 with a mask 678, where white represents the input and black represents the recomposed pixels. The token-based model 679 may generate a mask of pixels in the missing areas, recompose the pixels in the missing areas of the panoramic canvas, and merge the mask with the recomposed pixels. The token-based model 679 outputs a high-resolution panoramic video 680 having 22 frames (i.e., having a threshold number of pixels).

[0070] Once the panoramic video is finalized, it may be provided to the user who captured the input panning video. In some embodiments, the panoramic video is displayed with an option to select a panoramic image from the panoramic video. For example, the panoramic video may be displayed as a series of panoramic images, and the user interface may include an option to isolate one or more of the panoramic images.

[0071] Exemplary Methods 9 shows a flowchart of an example method 900 for generating a panoramic video. Method 900 may be performed by computing device 200 of FIG. 2. In some embodiments, method 900 is performed by user device 115, media server 101, or partially by user device 115 and partially by media server 101.

[0072] The method 900 of FIG. 9 may begin at block 902. At block 902, an input panning video is received that includes image frames captured by a camera that pans from a first side to a second side at least once. In some embodiments, the quality of the panning video improves as a function of the number of times the image frames captured by the camera include panning from the first side to the second side and panning from the second side to the first side. In some embodiments, the input panning video includes one or more moving objects. Block 902 may be followed by block 904.

[0073] In block 904, the input panning video is reprojected onto a panoramic canvas, which includes the image frames and the missing regions. Block 904 may be followed by block 906.

[0074] In block 906, the generative machine learning model outputs a panoramic video including synthesized pixels of the missing regions of the panoramic canvas by generating a base video including pixels of the missing regions at a coarse temporal scale, performing temporal upsampling, merging the temporally upsampled base video with the panoramic canvas, and refining the base video by resynthesizing the synthesized pixels of the missing regions.

[0075] Temporal upsampling may include temporally upsampling the base video and performing spatial and color alignment of the temporally upsampled base video with the panoramic canvas to obtain an aligned temporally upsampled base video. Merging the temporally upsampled base video with the panoramic canvas may include merging the aligned temporally upsampled base video with the panoramic canvas. Resynthesizing may include generating a mask of pixels of the missing region, regenerating synthesized pixels of the missing region based on the mask, and merging the regenerated pixels with the synthesized pixels.

[0076] In some embodiments, in response to outputting the panoramic video, a super-resolution pass is applied to the panoramic video to increase the resolution of the panoramic video. In some embodiments, a panoramic image is selected from the panoramic video.

[0077] In the above description, for purposes of explanation, numerous specific details are set forth to provide a thorough understanding of the specification. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In some instances, structures and devices are shown in block diagram form in order to avoid obscuring the specification. For example, embodiments may be described above primarily with reference to user interfaces and specific hardware. However, embodiments may apply to any type of computing device capable of receiving data and commands, and any peripheral device that provides services.

[0078] In addition to the above, users may be provided with controls that allow them to make choices about both whether and when the systems, programs, or features described herein may enable the collection of user information (e.g., information about the user's social network, social behavior, or activities, occupation, user preferences, or the user's current location), as well as whether the user is sent content or communications from the server. Furthermore, certain data may be processed in one or more ways so that personally identifiable information is removed before it is stored or used. For example, the user's identity may be processed so that personally identifiable information about the user cannot be determined, or if location information is obtained (e.g., to the city, zip code, or state level), the user's geographic location may be generalized so that the user's specific location cannot be determined. Thus, users may control what information is collected about them, how that information is used, and what information is provided to them.

[0079] A reference herein to "some embodiments" or "some instances" means that a particular feature, structure, or characteristic described in connection with an embodiment or instance may be included in at least one implementation herein. The appearances of the phrase "in some embodiments" in various places in this specification are not necessarily all referring to the same embodiments.

[0080] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. These steps require physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It is sometimes convenient, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0081] It should be recognized, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise specified, as will be apparent from the discussion that follows, throughout this specification, discussions utilizing terms including "processing," "calculating," "figuring out," "determining," or "displaying," etc., will be understood to refer to the actions and processes of a computer system or similar electronic computing device that manipulates and converts data represented as physical (electronic) quantities in the computer system's registers and memory into other data that is also represented as physical quantities in the computer system's memory or registers, or other such information storage, transmission, or display device.

[0082]

[0013] Embodiments herein may also relate to a processor for performing one or more steps of the above-described methods. The processor may be a special-purpose processor selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored on a non-transitory computer-readable storage medium, including, but not limited to, any type of disk, including an optical disk, a ROM, a CD-ROM, a magnetic disk, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory, including a USB key with non-volatile memory, or any type of medium suitable for storing electronic instructions, each connected to a computer system bus.

[0083] The specification may take the form of some entirely hardware embodiments, some entirely software embodiments, or some embodiments containing both hardware and software elements, hi some embodiments, the specification is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.

[0084] Furthermore, the specification may take the form of a computer program product accessible from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For purposes of this description, a computer-usable or computer-readable medium may be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, instruction execution apparatus, or instruction execution device.

[0085] A data processing system suitable for storing or executing program code will include at least one processor coupled directly or indirectly to memory elements via a system bus. The memory elements may include local memory utilized during the actual execution of the program code, bulk storage, and cache memory that provides temporary storage of at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution.

Claims

1. 1. A computer-implemented method comprising: receiving an input panning video comprising image frames captured by a camera panning at least once from a first side to a second side; and reprojecting the input panning video onto a panoramic canvas, the panoramic canvas including the image frames and the missing region, the computer-implemented method further comprising: and outputting a panoramic video including synthesized pixels of the missing region of the panoramic canvas using a generative machine learning model, the generative machine learning model comprising: generating a base video including the synthesized pixels of the missing region at a coarse time scale; performing temporal upsampling, merging the temporally upsampled base video with input pixels from the panoramic canvas, and refining the base video by recomposing the synthesized pixels of the missing regions; and outputting the panoramic video.

2. The temporal upsampling Temporally upsampling the base video; and performing spatial and color alignment of the temporally upsampled base video with the panoramic canvas to obtain an aligned temporally upsampled base video; 2. The computer-implemented method of claim 1, comprising:

3. 3. The computer-implemented method of claim 2, wherein merging the temporally upsampled base video and the panoramic canvas comprises merging the aligned temporally upsampled base video with the panoramic canvas.

4. 4. The computer-implemented method of claim 3, wherein the recombining comprises generating a mask of pixels of the missing region, regenerating the synthesized pixels of the missing region based on the mask, and merging the regenerated pixels with the synthesized pixels.

5. 2. The computer-implemented method of claim 1, wherein the quality of the panoramic video improves as a function of the number of times the image frames captured by the camera include panning from the first side to the second side and panning from the second side to the first side.

6. The computer-implemented method of claim 1 , wherein the input panning video includes one or more moving objects.

7. The computer-implemented method of claim 1 , further comprising, in response to outputting the panoramic video, applying a super-resolution pass to the panoramic video to increase a resolution of the panoramic video.

8. The computer-implemented method of claim 1 , further comprising selecting a panoramic image from the panoramic video.

9. A non-transitory computer-readable medium storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations, the operations including: receiving an input panning video comprising image frames captured by a camera panning at least once from a first side to a second side; reprojecting the input panning video onto a panoramic canvas, the panoramic canvas including the image frames and missing regions, the operations further comprising: and outputting a panoramic video including synthesized pixels of the missing region of the panoramic canvas using a generative machine learning model, the generative machine learning model comprising: generating a base video including the synthesized pixels of the missing region at a coarse time scale; performing temporal upsampling, merging the temporally upsampled base video with input pixels from the panoramic canvas, and refining the base video by recomposing the synthesized pixels of the missing regions; a non-transitory computer-readable medium for outputting the panoramic video by

10. The temporal upsampling Temporally upsampling the base video; and performing spatial and color alignment of the temporally upsampled base video with the panoramic canvas to obtain an aligned temporally upsampled base video; 10. The non-transitory computer-readable medium of claim 9, comprising:

11. 11. The non-transitory computer-readable medium of claim 10, wherein the merging the temporally upsampled base video and the panoramic canvas comprises merging the aligned temporally upsampled base video with the panoramic canvas.

12. 12. The non-transitory computer-readable medium of claim 11, wherein the recombining comprises: generating a mask of pixels of the missing region; regenerating the synthesized pixels of the missing region based on the mask; and merging the regenerated pixels with the synthesized pixels.

13. The non-transitory computer-readable medium of claim 9 , wherein the input panning video includes one or more moving objects.

14. 10. The non-transitory computer-readable medium of claim 9, wherein the operations further include, in response to outputting the panoramic video, applying a super-resolution pass to the panoramic video to increase a resolution of the panoramic video.

15. The non-transitory computer-readable medium of claim 9 , wherein the actions further include selecting a panoramic image from the panoramic video.

16. 1. A system comprising: a processor; a memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to perform operations, the operations including: receiving an input panning video comprising image frames captured by a camera panning at least once from a first side to a second side; reprojecting the input panning video onto a panoramic canvas, the panoramic canvas including the image frames and missing regions, the operations further comprising: and outputting a panoramic video including pixels composited into the missing region of the panoramic canvas using a generative machine learning model, the generative machine learning model generating a base video including the synthesized pixels of the missing region at a coarse time scale; performing temporal upsampling, merging the temporally upsampled base video with input pixels from the panoramic canvas, and refining the base video by recomposing the synthesized pixels of the missing regions; and outputting the panoramic video.

17. The temporal upsampling Temporally upsampling the base video; and performing spatial and color alignment of the temporally upsampled base video with the panoramic canvas to obtain an aligned temporally upsampled base video; 17. The system of claim 16, comprising:

18. 20. The system of claim 17, wherein the merging the temporally upsampled base video and the panoramic canvas comprises merging the aligned temporally upsampled base video with the panoramic canvas.

19. 20. The system of claim 18, wherein the recombining comprises generating a mask of pixels of the missing region, regenerating the synthesized pixels of the missing region based on the mask, and merging the regenerated pixels with the synthesized pixels.

20. The system of claim 16 , wherein the input panning video includes one or more moving objects.

Citation Information

Patent Citations

  • In-camera generation of high-quality composite panoramic image

    JP2010252312A

  • Image processing device and method, and program

    JP2011082918A

  • Image processing device, image processing method and program

    JP2013077228A

  • Image processing device, image processing method, and program

    JP2019079114A

  • Generation device and computer program

    JP2020136884A