Video frame interpolation assisted by fast auxiliary camera

By using a high-speed auxiliary camera to guide video frame interpolation, the method effectively addresses the challenge of increasing frame rate and improving video quality in complex motion scenes, achieving enhanced interpolation results.

WO2026096516A1PCT designated stage Publication Date: 2026-05-07DOLBY LABORATORIES LICENSING CORP
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2025-10-28
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Conventional video frame interpolation methods struggle with increasing the frame rate of low-frame-rate videos, especially when the motion in the scene is complex, such as with nonlinear motion, and they often rely on assumptions that may not hold in such scenarios.

Method used

Utilizing a high-speed auxiliary camera, like a SPAD camera, to capture high-frame-rate videos that guide the interpolation of low-frame-rate CMOS camera videos, employing a neural network architecture for synthesis and optical flow refinement to enhance the interpolation process.

Benefits of technology

The method achieves significant framerate upsampling, improving the quality and clarity of interpolated frames, particularly in scenes with complex motion, by leveraging the high-speed frames from the auxiliary camera to guide the interpolation of CMOS camera frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025052925_07052026_PF_FP_ABST
    Figure US2025052925_07052026_PF_FP_ABST
Patent Text Reader

Abstract

An example video frame interpolation method includes receiving a sequence of first frames captured by a first camera at a first framerate and a sequence of second frames captured by a second camera at a greater second framerate. The method further includes, for a pair of consecutive first frames, performing a first interpolation to obtain a first set of interpolated frames at the second framerate, the first interpolation being based on the pair of consecutive first frames and further based on a corresponding set of the second frames; performing a different second interpolation to obtain a second set of interpolated frames at the second framerate, the second interpolation being based on the pair of consecutive first frames and further based on the corresponding set of the second frames; and generating a third set of interpolated frames at the second framerate by fusing the first and second sets of interpolated frames.
Need to check novelty before this filing date? Find Prior Art

Description

VIDEO FRAME INTERPOLATION ASSISTED BY FAST AUXILIARY CAMERA 1. Cross-Reference to Related Applications

[0001] This application claims the benefit of priority from U. S. Provisional Patent Application No. 63 / 714,445 filed on October 31, 2024 and U. S. Provisional Patent Application No. 63 / 747,674 filed on January 21, 2025. and EP Application No. 25154438.3 filed on January 28, 2025, each of which is incorporated by reference herein in their entirety.2. Field of the Disclosure

[0002] Various example embodiments relate to frame rate upscaling and, more specifically but not exclusively, to video frame interpolation assisted by a fast auxiliary' camera.3. Background

[0003] Frame interpolation is the process of synthesizing in-between images from a given sequence of images. One application of this technique includes temporal upsampling to increase the refresh rate of videos or to create slow motion effects. For example, with digital cameras and smartphones, a user may capture multiple images within a short period of time. In some examples, interpolating between those images can lead to engaging videos that reveal more details of the scene motion, thereby delivering an even more pleasing sense of the moment than the original images.SUMMARY

[0004] Disclosed herein are various embodiments of video frame interpolation assisted by a fast auxiliary camera. In some examples, the main camera is a conventional color (e.g., RGB) CMOS camera, and the fast auxiliary camera is an event camera or a single photon avalanche diode array (SPAD) camera. In some examples, the frame rate of the fast auxiliary camera is greater than the frame rate of the main camera by a factor of ten or more. Of course, as will be understood and appreciated by the skilled person, the main camera and / or the fast auxiliary camera may be implemented in any other suitable form.

[0005] In one example, a video interpolation apparatus comprises at least one processor and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: receive a sequence of first frames representing a scene as captured by a first camera at a first frame rate;receive a sequence of second frames representing the scene as captured by a second camera at a second frame rate that is greater than the first frame rate; and for a pair of consecutive first frames, perform a first interpolation to obtain a first set of interpolated frames at the second frame rate, the first interpolation being based on the pair of consecutive first frames and further based on a corresponding set of the second frames; perform a different second interpolation to obtain a second set of interpolated frames at the second frame rate, the different second interpolation being based on the pair of consecutive first frames and further based on the corresponding set of the second frames; and generate a third set of interpolated frames at the second frame rate by fusing the first and second sets of interpolated frames.

[0006] In another example, a video interpolation method comprises: receiving a sequence of first frames representing a scene as captured by a first camera at a first frame rate; receiving a sequence of second frames representing the scene as captured by a second camera at a second frame rate that is greater than the first frame rate; and for a pair of consecutive first frames, performing a first interpolation to obtain a first set of interpolated frames at the second frame rate, the first interpolation being based on the pair of consecutive first frames and further based on a corresponding set of the second frames; performing a different second interpolation to obtain a second set of interpolated frames at the second frame rate, the different second interpolation being based on the pair of consecutive first frames and further based on the corresponding set of the second frames; and generating a third set of interpolated frames at the second frame rate by fusing the first and second sets of interpolated frames.

[0007] According to yet another example provided is a non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the above method.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which:

[0009] FIGS. 1A-1C are schematic views illustrating an electronic device with which various embodiments can be practiced.

[0010] FIGS. 2A-2B illustrate example binary frames captured by the SPAD camera of the electronic device shown in FIGS. 1 A-1C according to some examples.

[0011] FIG. 3 is a block diagram illustrating several SPAD camera virtual exposure frames formed using a sequence of SPAD frames according to some examples.

[0012] FIGS. 4A-4F pictorially illustrate virtual exposure frames corresponding to different values of N according to some examples.

[0013] FIG. 5 is a block diagram illustrating two inputs and a corresponding output generated using a frame interpolation method according to some examples.

[0014] FIG. 6 is a high-level block diagram illustrating a video interpolation pipeline that can be used to implement the method illustrated in FIG. 5 according to some examples.

[0015] FIG. 7 is a block diagram illustrating an Enhanced Multi-Frame Synthesis Interpolation (EMFSI) module used in the video interpolation pipeline of FIG. 6 according to some examples.

[0016] FIG. 8 is a block diagram illustrating the notations used in the description of the EMFSI module illustrated in FIG. 7.

[0017] FIG. 9 is a block diagram illustrating a single frame synthesis (SFS) block used in the EMFSI module of FIG. 7 according to some examples.

[0018] FIG. 10 is a block diagram illustrating a modified U-Net synthesis block used in the SFS block of FIG. 9 according to some examples.

[0019] FIG. 11 is a block diagram illustrating a feature alignment block used in the EMFSI module of FIG. 7 according to some examples.

[0020] FIG. 12 pictorially illustrates frame-interpolation improvements that can be achieved using the EMFSI module of FIG. 7 according to some examples.

[0021] FIG. 13 pictorially compares the frame-interpolation results of several different implementations (versions) of the EMFSI module of FIG. 7 according to some examples.

[0022] FIG. 14 is a block diagram illustrating the notations used in the description of the video interpolation pipeline of FIG. 6 according to some examples.

[0023] FIGS. 15A-15C provide a block diagram of a PWC-Net and pictorially illustrate the effect of noise on different levels thereof according to some examples.

[0024] FIG. 16 is a block diagram illustrating the multi frame optical flow refinement (MFOFR) module that can be used in the video interpolation pipeline of FIG. 6 according to some examples.

[0025] FIGS. 17A-17B pictorially illustrate improvements obtained with the MFOFR module of FIG. 16 according to some examples.

[0026] FIG. 18 is a block diagram illustrating the video interpolation pipeline of FIG. 6 according to some additional examples.

[0027] FIG. 19 is a block diagram illustrating a merge module used in the video interpolation pipeline of FIG. 6 and FIG. 18 according to some examples.

[0028] FIG. 20 is a block diagram illustrating a merge module used in the video interpolation pipeline of FIG. 6 and FIG. 18 according to some additional examples.

[0029] FIG. 21 is a flowchart illustrating a video frame interpolation method that can be used to implement the video interpolation pipeline of FIG. 6 and FIG. 18 according to some examples.

[0030] FIG. 22 is a block diagram illustrating a computing device used to implement the method of FIG. 21 according to some examples.DETAILED DESCRIPTION

[0031] FIGS. 1A-1C are schematic views illustrating an electronic device 100 with which various embodiments can be practiced. More specifically, FIG. 1A shows an overall perspective view of the electronic device 100. FIG. 1B shows a perspective view of a camera module 130 used in the electronic device 100. FIG. 1C shows a side view of the camera module 130. In the example shown, the electronic device 100 is a smartphone. In other examples, the camera module 130 may be a part of another suitable electronic device.

[0032] In the example shown, the camera module 130 of the electronic device 100 includes a first camera 110, e.g., a CMOS camera, and a second camera 120, e.g., a SPAD camera. Of course, as will be understood and appreciated by the skilled person, any other suitable camera may be used for the first camera 110 and / or the second camera 120. depending on various implementations and / or circumstances. Each of the cameras 110, 120 includes a respective optical lens barrel, a respective shutter mechanism, and a respective pixelated image sensor. The optical lens barrel typically includes a lens for collecting the light received from a scene anddirecting the collected light to the image sensor. The optical lens barrel may also include an optical mechanism for performing focus adjustment, changing optical magnification, and adjusting an amount of light directed to the image sensor.

[0033] The camera module 130 also includes a rectangular plate 140 configured to provide structural support to the cameras 110, 120 in the body of the electronic device 100 and further configured to fix the relative positions of the cameras 110, 120 with respect to one another. The body of the electronic device 100 also contains electronic circuits, such as an electronic controller and one or more image processors, to which the cameras 110, 120 are electrically connected via electrical connectors 112 and 122, respectively (also see FIG. 22). In at least some examples, the electronic device 100 includes a display, on which the images and videos captured with the cameras 110, 120 can be displayed.

[0034] Many things in life happen in the blink of an eye. To watch those events in slow motion, a high-speed video capture feature of the electronic device 100 may be used.Conventional CMOS cameras (e.g., the first camera 110) typically capture videos at a frame rate of 20-60 FPS. Some models of the iPhone may capture slow motion videos at 240 FPS. For comparison, SPAD cameras (such as the second camera 120), leveraging the corresponding single-photon imaging sensors, are able to capture binary frames (bit depth = 1) at up to ca. 100,000 FPS.

[0035] FIGS. 2A-2B illustrate example binary frames 202, 204 captured by the SPAD camera 120 according to some examples. More specifically, the frame 202 is captured under relatively bright illumination conditions. The frame 204 is similarly captured under relatively dark illumination conditions. As illustrated by FIGS. 2A-2B, in some cases, the video frames captured by SPAD cameras may be rather quantized and corrupted by Poisson noise, e.g., due to the photon counting mechanism of the SPAD sensor and the relatively short exposure time.

[0036] We have realized that the relatively low-quality high-frame rate video (e.g., captured by the exemplary SPAD camera 120, or the like) can beneficially be used to guide video frame interpolation of a high-quality low-framerate video (e.g., captured by the CMOS camera 110, or the like). Disclosed herein below is an example video interpolation pipeline configured to utilize the noisy high-framerate SPAD video to increase the frame rate of a clean but low-framerate CMOS camera video such that the user can obtain a video that has both high quality images and a high framerate. In particular, some examples may provide 100x upsampling in the framerate on example datasets.

[0037] Various example embodiments may be characterized by one or more of the following features:1. Use of the SPAD camera 120 to increase the frame rate of the content captured with the CMOS camera 110. Some examples provide a 1: 100 ratio in terms of the temporal frequency up-sampling.2. Neural network architecture (e.g., a synthesis module) configured to perform video interpolation using inputs from Virtual Exposure stacks created from the raw SPAD camera data and CMOS key frames.3. In some examples, the neural network architecture incorporates residual learning, feature alignment through deformable convolution, and multi-frame merging to improve the quality of an interpolated video frame.4. Additional architectures configured to leverage a motion-based mechanism (e.g., a warping module) to combine the inputs derived from SPAD and CMOS data for video frame interpolation. Some embodiments implement a mechanism configured to perform optical flow refinement using the optical flow obtained from the noisy SPAD frames. 5. In some embodiments, an overall architecture is configured to combine the information from the synthesis and warping modules to perform more effective frame interpolation.

[0038] The SPAD camera 120 operates to image a given scene by counting photons. Photon arrival at the SPAD sensor is a Poisson process and follows the Poisson distribution described by Eq. (1):A.ke~P X = k) = —— (1)K Iwhere X is a random variable representing the number of photons incident on the SPAD sensor; A is the expected number of photons arriving at the SPAD sensor during the course of exposure; and k is a positive integer representing the number of photons.

[0039] In some examples, the SPAD camera 120 is configured to image the scene at 100,000 binary' frames per second (BFPS) and has a dark count rate of 27 counts per second (cps). Such SPAD cameras are commercially available and include, for example, the Pi Imaging SPAD 5122camera.

[0040] For each SPAD binary' frame B, the pixel value B(x, y) = 1 if one or more photons arrived at the pixel location (x. y) during the exposure time of the binary frame, and B x, y) = 0 otherwise. Together with the photon arrival probability given in equation (1), the value of each pixel (x,y) in a binary frame B can be drawn from the Bernoulli distribution:where Bx yG {0, 1} is a Bernoulli random variable representing the value of pixel (x, y) in the binary frame B, <pxy(cps is the photon flux at pixel location (x. y), p (%) is the quantum efficiency, r (cps) is the dark count rate, and t sec) is the exposure time.

[0041] In some examples, due to the current relatively high cost and limited availability of SPAD camera hardware, we simulate the SPAD video data from the THU-HSEVI dataset, which features 8-bit monochrome 5000 FPS videos captured using a CMOS sensor. The THU-HSEVI dataset has six scenes of high-speed phenomena, such as smashing a cup or bursting a balloon, and the like. Each scene contains five independent videos.

[0042] In some examples, we simulate SPAD images from 8-bit RGB images. Given an 8-bit CMOS frame, we may not have access to the ground truth photon flux information. We target an average pixel detection rate of ~0.1 photons per pixel per binary frame (ppbf) across all frames in a given video for this simulation, which represents a low photon flux scenario, e.g., in line with the short exposure time (10‘5sec) of each binary frame. Given a CMOS frame FT, where T G [0, 1, 2,, n^rames} is the frame at time stamp T in a given video, and nframesis the total number of frames in the video, we generate the SPAD binary frame BTat time stamp T based on Eqs. (3)-(5):where FTI(x,y) is the pixel value at location (x, y) of frame FT, avg pixel value is the average pixel value across all CMOS frames in a given video, dimxis the x dimension (number of columns) of the frame FT, dimyis the y dimension (number of rows) of the frame FT,'where rv = 2.7 e — 4 is the expected number of dark counts given a dark count rate r of 27 cps and an exposure time T of IxlO'5sec (IxlO5FPS), and y = 9 is a hyperparameter chosen to meet the target average cps per pixel per frame value, andBT(x,y) = Bxy(5)where Bx yG {0, 1} is a Bernoulli random variable representing the value of pixel (x, y) in the binary frame BT, and BT(x, y) is the pixel value at location (x, y) of the binary’ frame BT.

[0043] Since in some examples, the SPAD camera 120 captures images at 105FPS, and the CMOS camera 110 captures images at 5000 FPS, at each time stamp T we generate 20 SPAD binary frames B1, B2,..., B20for each CMOS frame FT. and use the average of the 20 binary frames to generate the corresponding SPAD frame STas follows:where STis an 8-bit integer array of dimension C x340x260 (unless otherwise specified, the dimensions in this specification are given in the order of CxHxW, where C=3 for RGB frames and C=1 for monochrome frames).

[0044] FIG. 3 is a block diagram illustrating several SPAD camera virtual exposure frames formed using a sequence 300 of SPAD frames according to some examples. Given a SPAD frame STat time T, one can create a Virtual Exposure (VE) frame VE of STby averaging its temporally nearest N frames as illustrated in FIG. 3. When N=l, VE^ represents the SPAD frame STitself. For N>1, the VE frame VE is obtained as follows:For illustration purposes, virtual exposure stacks 302, 304, 308, and 316 corresponding to N = 2.4, 8, and 16, respectively, are indicated in FIG. 3. By choosing different N values, one can obtain different VE frames that represent different levels of noise-to-blur tradeoff. For example, the higher the value of N, the lower the level of noise and the higher the level of blur. The value of N is a hyperparameter, the suitable value of which can be determined experimentally. In setting this value, one may generally take into consideration the amount of motion and the amount of noise in the corresponding dataset.

[0045] FIGS. 4A-4F pictorially illustrate VE frames 402-412 corresponding to different values of N according to some examples. More specifically, the VE frames 402-412 correspond to N = 2, 4, 8, 16, 32, and 64, respectively. As evident from FIGS. 4A-4F, the broken cup fragment (located within the in-frame rectangle) is significantly obscured by the noise for VE2rrepresented by the frame 402 (FIG. 4A). As the value of N increases, the noise level decreases, and the cup fragment becomes more clearly discernible. However, as the value of N keeps increasing towards 64, the blurriness starts to increase as a result of motion averaging multiple misaligned frames, e.g., as can be noticed for VE^ represented by the frame 412 (FIG. 4F).

[0046] In one example, a VE stack VTis a collection of a range of VEs at time T. In some embodiments, we use a VE stack that incorporates a set of VEs with the N values of 1, 2. 4, 8, 16, 32, and 64. Mathematically, such VE stack VTis represented as:where VTis a 7x340x260 array, VE is a C X340X260 array, and cat represents a concatenation operation along the channel dimension.

[0047] In at least some examples, SPAD cameras have several desirable properties leading to many applications in imaging and photography. First, SPAD cameras are able to capture videos at a significantly higher framerate compared to that of t pical CMOS cameras. This characteristic allows for the imaging of high-speed scenes and phenomena, such as the trajectory of a pulse of light. Since a SPAD camera images a scene by counting photons, in theory, saturation is not an issue, and the camera may be considered to have a substantially infinite dynamic range. This characteristic allows for very High Dynamic Range (HDR) photography in just a single capture, without using any exposure bracketing techniques. SPAD cameras are also very sensitive to light, which makes them suitable for low light imaging. Additionally, SPAD cameras may be advantageous for some medical imaging applications, such as imaging very weak fluorescence signals used, for example, to guide disease diagnostics and surgical procedures.

[0048] Under some conventional approaches, video frame interpolation is performed using the low frame-rate video itself, without the help of an auxiliary camera. Such approaches may rely on certain assumptions with regard to the motion trajectories between the existing frames, which may be challenging when the motion in the scene is complicated, such as when nonlinear motion is present.

[0049] In some cases, the use of an auxiliary camera or sensor can bring additional useful information for video frame interpolation. For example, a fast auxiliary' camera, such as a highspeed camera or an event camera, can be used to fdl in the missing motion between the corresponding neighboring frames of the main (slow) camera. Accordingly, some embodiments are directed to using the high-speed frames from the SPAD camera 120 to guide video interpolation of the frames from the CMOS camera 110.

[0050] FIG. 5 is a block diagram illustrating two inputs 502, 504 and a corresponding output 506 generated using a frame interpolation method 500 according to some examples. For illustration purposes and without any’ implied limitations, the input 502 includes frames FTcaptured at 50 FPS with the CMOS camera 110. The input 504 includes frames ST captured at 5000 FPS with the SPAD camera 120. As such, in the example shown, the frame-rate ratio for the inputs 502, 504 is one hundred. In other examples, other frame-rate ratio for the inputs 502, 504 can also be used. As illustrated in FIG. 5, given the low frame rate CMOS video frames FTand high framerate SPAD video frames ST, the method 500 generates a corresponding sequence of interpolated frames FTfor the output 506. In the example shown, the output 506 has the same frame rate as the input 504.

[0051] FIG. 6 is a high-level block diagram illustrating a video interpolation pipeline 600 that can be used to implement the method 500 according to some examples. The pipeline 600 includes a warping interpolation module 610 and an Enhanced Multi-Frame Synthesis Interpolation (EMFSI) module 620. The pipeline 600 further includes an attention-based merge module 630 configured to fuse interpolation results 612, 622 received from the interpolation modules 610. 620 to generate the output 506 as illustrated in FIG. 5. The frame notations used in the following description of the pipeline 600 are also indicated in FIG. 5.

[0052] FIG. 7 is a block diagram illustrating the EMFSI module 620 used in the video interpolation pipeline 600 according to some examples. Given the SPAD Virtual Exposure (VE) stacks VT-1, VT, VT+1at time T — 1, T, T + 1, and their respective temporally nearest CMOS key framesFT, FT+1, one function of the EMFSI module 620 is to reconstruct an image with features that are temporally aligned with the fast SPAD frame ST at time T and an appearance that is consistent with its immediately preceding and succeeding CMOS key frames. In the example shown, the EMFSI module 620 includes three stages: a single frame synthesis stage 710, a feature alignment stage 720, and a multi-frame enhancement stage 730.

[0053] FIG. 8 is a block diagram 800 illustrating the notations used in the description of the EMFSI module 620 shown in FIG. 7. These notations are used below in the description of FIGS.7 and 9-1 1 illustrating various components of the EMFSI module 620 according to some examples.

[0054] Referring to FIGS. 6-8, an output 622 of the EMFSI module 620 is denoted as FFMFSIand is expressed as follows:where EMFSI is the effective transfer function of the EMFSI module 620, and Fey(C X340X260) is the temporally nearest CMOS key frame with respect to the SPAD frame ST at time T. In one example, F^eyis obtained based on Eq. (10):where a = 1 for mod(T, 100) < 50, and a = 0 otherwise. This definition assumes the video frame interpolation factor of 100. FFrevis the adjacent CMOS key frame immediately preceding the SPAD frame ST at time T, and F"extis the adjacent CMOS key frame immediately following the SPAD frame ST at time T. FFMFSI(C X340X260) denotes the interpolated frame 622 at time T.

[0055] FIG. 9 is a block diagram illustrating a single frame synthesis (SFS) block 714 used in the EMFSI module 620 according to some examples. More specifically, in the example shown in FIG. 7, the stage 710 of the EMFSI module 620 employs three instances of the SFS block 714, with each instance being configured to receive a different respective pair of inputs 1, 2. In the configuration shown in FIG. 9, the input 1 includes the VE stack VT. and the input 2 includes the key frame Fey. Other configurations of the SFS block 714 used in the stage 710 of the EMFSI module 620 are indicated in FIG. 7.

[0056] During the single frame synthesis stage 710, the input pairs including the SPAD VE stack VTand the nearest CMOS key frame Feyat times T — 1, T, and T + 1 ((Vr-i, FFly),are passed into the respective instances of the SFS block 714 to obtain an intermediate synthesis frame FFFSat their respective time stamps. The corresponding functionality is mathematically expressed as follows:where fSFSrepresents the transfer function of the SFS block 714, and FFFS(C X340X260) is the intermediate synthesized frame from the SFS block for the time T.

[0057] As indicated in FIG. 9, a first part of the SFS block 714 includes a modified U-Net synthesis block 910 comprising an encoder-decoder structure with skip connections therebetween. The modified U-Net synthesis block is constructed by modifying a conventional U-Net architecture. One modification of the conventional U-Net structure used to arrive at the modified U-Net synthesis block 910 is to use a two-layer ResNet in the encoder part of the U-Net structure instead of the conventionally used double convolution. The encoder part of the U-Net synthesis block 910 takes, as inputs 1, 2, the SPAD virtual exposure stack VTand the nearestCMOS key frame Fey, respectively, and the decoder part of the U-Net synthesis block 910 outputs a synthesized frame fflietRes(C x340x260).

[0058] FIG. 10 is a block diagram illustrating the modified U-Net synthesis block 910 used in the SFS block 714 according to some examples. The corresponding functionality is mathematically expressed as follows:where fUnetRes represents the transfer function of the modified U-Net synthesis block 910. In the example shown, the modified U-Net synthesis block 910 comprises a ResNet block 1002, the structure of which is illustrated in the corresponding first expansion inset. The modified U-Net synthesis block 910 further comprises five instances of a downsample block 1004 of varying sizes. An example structure of the downsample block 1004 is illustrated in the corresponding second expansion inset. The modified U-Net synthesis block 910 further comprises five instances of an upsample block 1006 of varying sizes. An example structure of the upsample block 1006 is illustrated in the corresponding third expansion inset.

[0059] Referring back to FIG. 9, the output ffnetResof the U-Net synthesis block 910 is refined in the SFS block 714 using a deformable-convolution-based approach. In the example shown, a convolution layer 920 is used to process the output ffnetRestogether with its nearest CMOS key frame Feyto generate the Offset (lx3x 3) and the Mask (lx3x 3), which serve as the two parameters taken by a deformable convolution block 930 to shape its deformable kernels. The corresponding functionality of the convolution layer 920 is mathematically expressed as follows:where Conv2D denotes a 2D convolution with kernel size = 3, stride = 1, padding = 1.

[0060] The deformable convolution block 930 then uses the Offset and Mask parameters received from the convolution layer 920 to apply deformable convolution on the nearest CMOS frame Feyto obtain Aey(C x340x 260), which is feature-aligned to the output ffnetResThe corresponding functionality of the deformable convolution block 930 is mathematically expressed as follows:A^ey= Def ormConvff^f Off set, Mask (15)where DeformConv denotes deformable convolution with kernel size = 3, stride = 1, padding = 1.

[0061] Thereafter, the SFS block 714 operates to calculate the difference between pfnetKesand Aeyand passes the calculated difference through a convolution layer 940. Finally, the output of the convolution layer 940 is added to pfnetRest0generate a final output FfFSof the SFS block 714. By performing these operations, the SFS block 714 operates to pass the aligned features over from the nearest key frame Feyand then uses this result to further refine the output pfnetResof the modified U-Net synthesis block 910. The corresponding functionality of the last portion of the SFS block 714 is mathematically expressed as follows:where Conv D is a 2D convolution with kernel = 3, stride = 1, padding = 1.

[0062] FIG. 11 is a block diagram illustrating a feature alignment block (FA block) 724 used in the stage 720 of the EMFSI module 620 according to some examples. More specifically, in the example shown in FIG. 7, the stage 720 of the EMFSI module 620 employs two instances of the FA block 724, with each instance being configured to receive a different respective pair of inputs, as indicated in FIG. 7. Given the intermediate synthesized frames FfFf FfFS, Fff received from the plurality of SFS blocks 714, one function of the feature alignment (FA) blocks 724 is to align the frames F, andfrom times T and F+l to the frame FfFSat time T. Similar to the refinement procedure used in the SFS block 714, the FA block 724 first uses a convolution layer 1110 to calculate the Offset and the Mask using the frames FfFSand FFS(T T') in accordance with Eqs. (13) and (14), respectively. The FA block 724 then uses a deformable convolution block 1120 to apply deformable convolution to FFSin accordance with Eq. (15) to obtain ASTFSthat is aligned to the frame FfFS. The FA block 724 subsequently operates to obtain the difference between FfFSand AFS. A convolution layer 1130 configured to operate in accordance with Eq. (16) is used to process the obtained difference, and the output of the convolution layer 1130 is then added to the frame FfFSto obtain the feature aligned output frame F^. For the first instance of the FA block 724 used in the stage 720, T'= 7-1. For the second instance of the FA block 724 used in the stage 720, T'= T+ 1.

[0063] Referring back to FIG. 7, in the example shown, the stage 730 of the EMFSI module 620 employs a modified U-Net merge block 734. In some examples, the modified U-Net merge block 734 is implemented using a U-Net architecture similar to that of the denoising block usedin FastDVDNet described elsewhere. The modified U-Net merge block 734 is configured to take, as inputs, the frames FT-I, T’ F'TFS. F^+I, Tar|dtooutput a final synthesized frame at FFMFSI(Cx340x260) corresponding to time T. The corresponding functionality of the modified U-Net merge block 734 is mathematically expressed as follows:where fMFMis the transfer function of the modified U-Net merge block 734.

[0064] In one example, the above-described architectures use a loss function that includes an LI loss between the synthesized frame FFMFSIand the ground truth frame FFT(Cx340x260), as well as the LI loss between the x and y direction gradients (Sobel Filter) of the frames FFMFS / and F-Ftto encourage a better reconstruction on the edge features. The mathematical expression of the corresponding loss function is given by Eq. (18):where V denotes the gradient operation. In some examples, the coefficient has a value of? = 0.1. In other examples, other (than LI) loss functions can effectively be used. Examples of such other loss functions include, but are not limited to, SSIM / L2 functions and functions including combination of different types of loss functions.

[0065] In some examples, the above-described architectures can be tested on a subset of the THU-HSEV dataset. More specifically, in one example, one can train and test the EMFSI module 620 on the “cup breaking”, “chessboard waving”, and “hard disc spinning” scenes of the THU-HSEV dataset. For example, the EMFSI module 620 can be trained on videos 1, 3, 5 of the three scenes respectively, and then tested on videos 2 and 4 of those three scenes, e g., as illustrated in more details below.

[0066] In the examples shown, we train the EMFSI module 620 for 115 epochs using the ADAM optimizer with an initial learning rate of 0.01 and a batch size of 16.

[0067] FIG. 12 pictorially illustrates frame-interpolation improvements that can be achieved using the above-described architectures for the EMFSI module 620 according to some examples. More specifically, FIG. 12 pictorially compares the video interpolation results obtained with the EMFSI module 620 and the video interpolation results from a state-of-the-art video interpolation algorithm EMA-VFI for the “cup breaking” scene. More specifically, the first (top) row of images shows the sequence of frames obtained using the EMA-VFI algorithm. The second row of images shows the sequence of frames obtained using the EMFSI module 620. The third row of images shows the sequence of ground truth (GT) frames. The fourth (bottom) row of imagesshows the sequence of SPAD frames. The qualitative results corresponding to FIG. 12 are listed in Table 1, which is provided below.

[0068] FIG. 13 pictorially compares the frame-interpolation results of several different implementations (versions) of the EMF SI module 620 according to some examples. The shown comparison can be used to illustrate the rationale behind the design decisions for some variants of the EMFSI module 620. The corresponding modifications of the EMFSI module include the following:1. Using only the Stage 710 Single Frame Synthesis.2. Using the Stage 710 Single Frame Synthesis and directly feed the outputs thereof to the stage 730 Multi-Frame Enhancement without going through the Stage 720 Feature Alignment.

[0069] The first (leftmost) column in FIG. 13 shows example images obtained via conventional single frame synthesis. The second column in FIG. 13 shows example images obtained using the above-indicated modification number 2. The third column in FIG. 13 shows example images obtained using the full version of the EMFSI module 620 illustrated in FIG. 7. The fourth (rightmost) column in FIG. 13 shows the corresponding ground truth images.

[0070] Based on the example visual and quantitative results presented in FIGS. 12-13 and Table 1, respectively, we find that the use of the multi-frame enhancement stage 730 and incorporation of the pertinent information from two adjacent frames FTand FT+I, T help to reduce or eliminate at least some of the artifacts around the edges of the interpolated frame. In some examples, the artifacts are likely due to the temporal misalignment of the nearest key CMOS frame and the SPAD frame at time T. The multi-frame enhancement stage 730 tends to improve the overall Structural Similarity (SSIM) value of the interpolated frame. However, naively passing the three frames F^i,T, F^FS. and F +rinto the MFM block tends to decrease the overall Peak Signal to Noise Ratio (PSNR) value of the interpolated frame. In some examples, to address these issues, we add an explicit feature alignment step before the MFM block. The test results show that doing so not only substantially restores the PSNR performance but also suppresses the artifacts around the edge regions even more, and further increases the SSIM value of the interpolated frame.Table 1: Quantitative Comparison

[0071] The visual results presented in FIG. 12 indicate that the EMFSI module 620 tends to perform better than the EMA-VFI algorithm for the two test scenes. For example, it can be seen in FIG. 12 that the EMA-VFI algorithm performs sufficiently well on interpolating the falling hammer in the scene. However, when it comes to reconstructing the broken cup at the intermediate time stamps, the EMA-VFI algorithm is not as good in general due to the large amount of non-linear motion and deformation present in that part of the scene. In contrast, the EMFSI module 620 is able to reconstruct the cup more faithfully in that part of the scene.

[0072] Another way to interpolate a new frame between two existing CMOS frames is by warping the CMOS key frame to the intermediate temporal location, given the optical flow between the key frame and the frame to be interpolated. The optical flow is estimated from the SPAD frames. Hereafter, this approach is referred to as warping-based frame interpolation (WBFI) and is implemented in the video interpolation pipeline 600 using the warping interpolation module 610 (also see FIG. 6).

[0073] Let us consider two CMOS key frames Ftand Ft+1at time stamps t and t + 1, and the objective is to interpolate a frame Ft+k / W0at time t + k / 100, where {k | k G Z, 0 < K < 100} represents the frame number within the 99 frames to be interpolated. In one example, Ft+k / ioocanbe obtained using Eq. (19):where W[^+k / wodenotes the warping operation of Ftfrom the time stamp t to the timestamp t + k— using FLt t |h. which is the optical flow between the CMOS frame Ftand the CMOS frame^t+fc / ioo- In oneexample, the coefficient a is a = 1 for mod(k, 100) < 50, and is a = 0 otherwise.

[0074] FIG. 14 is a block diagram 1400 illustrating the notations used in the description of the WBFI approach in general and the warping interpolation module 610 in particular, according to some examples.k

[0075] Note that the CMOS frame Ft+k / 100at the intermediate timestamp t + — does not exist. Directly estimating the optical flow FLk using the CMOS frames Ftand F k does not appear feasible. To address this issue, we consider leveraging the fast SPAD frames Stl.1 2 99which exist everywhere at time stamps t (vt E {1t, t H - 100, t H - 100,, t H - 100, t + 1}J) between the key CMOS frames Ftand Ft+1using Eq. (20). Calculating the backward flow from frame Ft+1to Ft+k / 100follows the same procedure as indicated in Eq. (21).is the optical flow from SPAD frame S to SPAD frame S +it+ — t+ —

[0076] Finding the inter-frame motion through Eqs. (20) and (21 ) includes estimating the optical flow between noisy SPAD video frames. However, the optical flow estimation between noisy frames is a lesser researched topic. Since the optical flow estimation is based on matching pixels with similar intensity values between two different frames, the presence of noise (such as the presence of abrupt intensity changes) may decrease the accuracy of the estimated optical flow. Some conventional approaches use patch-wise intensity matching methods, e.g., analogous to applying a low-pass filter, to counter the effect of noise. However, a patch-based matching method can only return a spatially coarse optical flow at the patch level, not a spatially dense optical flow at the pixel level. In contrast, example embodiments disclosed herein use a learning-based optical flow estimation module configured to use coarse-to-fine flow estimation together with flow refinement based on the temporal consistency of the optical flow between neighboring frames. Hereafter, this module is referred to as the Multi Frame Optical Flow Refinement (MFOFR) module.

[0077] FIGS. 15A-15C provide a block diagram of a PWC-Net 1500 and pictorially illustrate the effect of noise on different levels of the PWC-Net 1500 according to some examples. In one example, the PWC-Net 1500 is constructed using a plurality of CNNs configured to determine the optical flow using the Pyramid, Warping, and Cost Volume operations. Under the above-mentioned WBFI approach, we first estimate the optical flow between two adjacent frames using a neural network constructed based on the PWC-Net 1500, the pyramid structure of which is illustrated in FIG. 15A. In the example shown, the PWC-Net 1500 estimates the optical flow between two frames in a spatially coarse-to-fine manner. This feature is designed to handle objects of various sizes and motions at various scales. However, we also find that this feature makes the motion estimations more robust to noise. In particular, we note that the effect of noise on the PWC-Net 1500 is small up until the second finest pyramid level, where feature matching is done on the patch-level. The error increases when the PWC-Net 1500 takes the coarse motion from the second finest pyramid level down to the finest pixel-level using a bilinear interpolation, as illustrated in FIGS. 15B and 15C.

[0078] FIG. 16 is a block diagram illustrating the MFOFR module 1600 that can be used in the warping interpolation module 610 according to some examples. Based on the above observations, the MFOFR module 1600 is designed to include a modified PWC-Net 1610, which is used for the initial coarse optical flow estimation between two SPAD frames. The modified PWC-Net is consistent with the original PWC-Net until the second finest pyramid level, and a learnable up-sampling layer is employed to replace the naive bilinear interpolation in the original PWC-Net 1500 to obtain a more accurate pixel-level optical flow estimation. In the example shown, the MFOFR module 1600 includes five instances of the modified PWC-Net 1610, with each of the instances being configured to process a different respective pair of SPAD frames. The temporal coherence of Optical Flow between adjacent frames is leveraged to regularize the Optical Flow estimation between the frames of interest by incorporating the Optical Flow from the neighboring frames using a tree-structured U-Net architecture designed by appropriately adapting the FastDVDnet. Accordingly, the MFOFR module 1600 also includes denoising blocks 1620, 1640 and is implemented using the adapted FastDVDnet. Note that the denoising blocks 1620, 1640 operate on different respective sizes of inputs. Three instances of a learnable upsampling layer 1630 are used to convert the outputs of the three instances of the denoising block 1620 to the size accepted by the denoising block 1640.

[0079] Given two SPAD frames S, i and Sfi+i, the MFOFR module 1600 is configured to use the modified PWC-Net 1610 to obtain an initial spatially-coarse Optical Flow estimationFLSPAi+1. The FLSPAPl+1is then passed together w ith the initial spatial ly-coarse OpticalFlow estimation from its neighboring frames FLSPADt;and FLSPAPt i+2into thet+ —, t+ — t+ —, t+ —FastDVDnet denoising block 1620 followed by the learnable upsampling layer 1630 configured to perform a 2D convolution, followed by a pixel shuffle layer and a bilinear interpolation to get to get back to the original resolution of the input image. As previously shown in FIG. 10, where the pertinent operations are illustrated on the RGB frames, these operations are instead performed for the neighboring intermediate Optical Flow triplets FLSPP_,,The resulting outputs are thensent into the final FastDVDnet denoising block 1640 to obtain the final optical flow estimation

[0080] In one example, we train our optical flow estimation network on the “Hard Disc Spinning,” “Chessboard Waving,” and “Balloon Bursting” scenes of the THU-HSEVI dataset. Since the dataset does not have the ground truth optical flow between the SPAD frames Sand St+t+i, we design our loss function to be based on the LI loss between the CMOS frame F i+i (which we have access to in the training data) and WZFL'Pi4Di+1(F _), which is Fwarped tow ards F i+i using the optical flow FLSPA?i+1. The mathematical expression of the corresponding loss function is given by Eq. (22):In this example, we train our model for 100 epochs using the ADAM optimizer with an initial learning rate of 0.01 and a batch size of 32.

[0081] FIGS. 17A-17B pictorially illustrate improvements obtained with the MFOFR module 1600 according to some examples. The shown examples are obtained using the above- mentioned “Hard Disc Spinning” dataset. Images 1702, 1704 shown in FIG. 17A are obtained by applying a warping operation configured using the refined optical flow determined using the MFOFR module 1600. In contrast, Images 1712, 1714 shown in FIG. 17B are obtained by applying a warping operation configured using the optical flow determined in a conventional manner. The presented results indicate that the MFOFR module 1600 is able to improve theaccuracy of Optical Flow between noisy frames in at least some occasions, such as the spinning disc scene illustrated in FIGS. 17A-17B.

[0082] In some cases, additional modifications can be made to the MFOFR module 1600 to improve video interpolation. First, we find that the MFOFR module 1600 trained on the aforementioned scenes of the THU-HSEVI dataset tends to zero out small motion on some scenes. Second, in some cases, error may accumulate as we sum up the motion in a manner consistent with Eqs. (20) and (21). In particular, since we may be interpolating a much larger number of frames compared to most other video interpolation tasks, in some examples, we may need to incorporate a flow refinement block to fix the motion in order for the motion estimation to be more suitable for the video interpolation task at hand.

[0083] FIG. 18 is a block diagram illustrating the video interpolation pipeline 600 according to some additional examples. In the example shown, the warping interpolation module 610 includes an optical flow estimation block 1802, an optical flow summation block 1804, an optical flow refinement block 1806, and a warping block 1808 serially connected to one another as indicated in FIG. 18. The optical flow refinement block 1806 implemented using the MFOFR module 1600 is placed after the flow summation block 1804 and before the warping block 1808 as illustrated in FIG. 18. In some examples, the optical flow refinement block 1806 is optional and may be omitted. As mentioned previously, the error accumulation that may occur when summing up the optical flow estimated between the adjacent SPAD frames is a reason why the warping-based interpolation module may have difficulties in some cases. The incorporation of the optical flow refinement block 1806 into the warping interpolation module 610 as indicated in FIG. 18 beneficially addresses this problem in such cases.

[0084] Some embodiments described above are configured to use attention-based modules, such as the deformable convolution, to obtain improved reconstruction results. In some alternative embodiments, other attention-based architectures, such as transformers, may also be used.

[0085] Due to the limitation of the THU-HSEVI dataset, which only offers monochrome video frames at 5000 FPS, further testing of the disclosed method on RGB data may be beneficial. The corresponding modifications are relatively straightforward and may include changing (increasing) the number of channels.

[0086] FIG. 19 is a block diagram illustrating the merge module 630 used in the video interpolation pipeline 600 according to some examples. As previously indicated, the mergemodule 630 is configured to fuse the interpolated frames 612, 622 received from the interpolation modules 610, 620, respectively, to generate the corresponding output frame 506. In one example, the merge module 630 is jointly trained with the other blocks / modules used in the pipeline 600.

[0087] In the example shown, copies of the interpolated frame 612 are applied to a concatenator (C) 1906 and a multiplier 1924. Copies of the interpolated frame 622 are similarly applied to the concatenator 1906 and a multiplier 1928. The concatenator 1906 operates to concatenate the interpolated frames 612, 622, thereby generating a concatenated frame 1908. A merge process block 1910 and a softmax function 1920 are then used as indicated in FIG. 19 to determine weighting coefficients wl and w2 based on the concatenated frame 1908. In one example, the merge process block 1910 is implemented using a U-Net architecture that takes as input the concatenated frame 1908 and outputs a tensor 1912 having the same resolution as the input frame. In another example, the merge process block 1910 can be implemented using two or more convolutional layers stacked one after the other. In some examples, the merge process block 1910 may be configured to use filters with a kernel size k x k (such as 3 x 3 or 5 x 5). The multipliers 1924, 1928 and an adder 1932 are configured to compute a weighted sum of the interpolated frames 612, 622, thereby generating the corresponding output frame 506.

[0088] FIG. 20 is a block diagram illustrating the merge module 630 used in the video interpolation pipeline 600 according to some additional examples. The embodiment illustrated in FIG. 20 differs from the embodiment of FIG. 18 in that the merge module 630 additionally includes encoders 2004, 2008 configured to generate feature maps 2006, 2010 representing the frames 612 and 622, respectively. The merge process block 1910 then operates on a corresponding concatenated feature map instead of operating on the frames 612, 622 directly. In one example, each of the encoders 2004, 2008 is implemented using a corresponding neural network configured to generate feature maps based on the corresponding image frames.

[0089] FIG. 21 is a flowchart illustrating a video interpolation method 2100 that can be used to implement the video interpolation pipeline 600 illustrated in FIG. 6 or FIG. 18 according to some examples. The method 2100 is described below with continued reference to FIGS. 1, 3, 5-7. 9-11.16, and 18-21.

[0090] A block 2102 of the method 2100 includes (i) receiving a sequence of first frames representing a scene as captured by a first camera at a first frame rate and (ii) receiving a sequence of second frames representing the scene as captured by a second camera at a secondframe rate that is greater than the first frame rate. In some examples, the first camera comprises a CMOS pixelated sensor (e.g., 110, FIGS. 1A-1C), and the first frames are RGB color frames. In some examples, the second camera comprises a single photon avalanche diode array (SPAD) sensor (e.g., 120, FIGS. 1 A-1C), and the second frames are binary frames. In some examples, the first and second frame rates differ by a factor of ten or more (e.g., by a factor of 100, as in FIG. 5).

[0091] A block 2104 of the method 2100 includes, for each time T of the sequence of second frames, generating a respective stack of virtual exposure frames (e.g., as illustrated in FIG. 3). Each of the virtual exposure frames is an average of N second frames nearest to the time T in the sequence. The respective stack of virtual exposure frames includes two or more virtual exposure frames corresponding to different respective values of N. In one example, the respective stack includes virtual exposure frames corresponding to N = 1, 2, 4, 8, 16, 32, and 64.

[0092] A block 2106 of the method 2100 includes, for a pair of consecutive first frames, performing a first interpolation to obtain a first set of interpolated frames at the second frame rate. The first interpolation is based on the pair of consecutive first frames and further based on a corresponding set of the second frames. In some examples, the first interpolation includes motion-based frame interpolation and is performed using the warping interpolation module 610 (e.g., see FIGS. 6 and 18). In some examples, the first interpolation is based on the respective stack corresponding to the time T. In some examples, the first interpolation is performed using a neural network comprising a plurality of PWC-Net blocks (e.g., 1610, FIG. 16), a plurality of denoising blocks (e.g., 1620, 1640, FIG. 16) connected to the plurality of PWC-Net blocks, and a plurality of upsampling layers (e.g., 1630, FIG. 16), with each of the upsampling layers being connected at least to a respective one of the denoising blocks.

[0093] In some examples, operations of the block 2106 include warping a key frame corresponding to the pair of consecutive first frames to an intermediate temporal position. In some examples, generating the key frame includes computing a weighted sum of first and second constituent frames of the pair. In some examples, the warping comprises: (i) estimating optical flow between the key frame and an interpolated frame at the intermediate temporal position based on optical flow between a corresponding pair of the second frames; and (ii) obtaining the interpolated frame at the intermediate temporal position by warping the key frame using the estimated optical flow. In some examples, the optical flow is estimated using the PWC-Net block 1610 (FIG. 16). In some examples, the estimating and obtaining operations are implemented using a neural network trained based on a loss function including an LI lossbetween a respective training frame at the intermediate temporal position and a corresponding one of the second frames warped towards the training frame using the optical flow (e.g., as expressed by Eq. (22)).

[0094] A block 2108 of the method 2100 includes, for a pair of consecutive first frames, performing a second interpolation to obtain a second set of interpolated frames at the second frame rate. In some examples, the second interpolation includes motion-independent frame interpolation and is performed using the EMFSI module 620 (e.g., see FIGS. 6 and 18). The second interpolation is based on the pair of consecutive first frames and further based on the corresponding set of the second frames. In some examples, the second interpolation for the time T is based on a plurality of the respective stacks corresponding to a set of frames adjacent to the time T.

[0095] In some examples, operations of the block 2108 include generating a key frame corresponding to the pair of consecutive first frames by computing a weighted sum of first and second constituent frames of the pair. The second interpolation for the time T is also based on the key frame. In some examples, for the time T, operations of the block 2108 also include (i) generating a first intermediate synthesis frame based on the respective stack of virtual exposure frames corresponding to time T-l: (ii) generating a second intermediate synthesis frame based on the respective stack of virtual exposure frames corresponding to the time T; and (iii) generating a third intermediate synthesis frame based on the respective stack of virtual exposure frames corresponding to time T+l. In some examples, the first, second, and third intermediate synthesis frames are generated using U-Net synthesis blocks (e.g., 910, FIG. 9). In some examples, generating the first, second, and third intermediate synthesis frames comprises deformable convolution (e.g., as expressed by Eqs. (13), (14)).

[0096] In some examples, operations of the block 2108 also include (a) generating a fourth intermediate synthesis frame via a feature alignment operation applied to the first and second intermediate synthesis frames; and (b) generating a fifth intermediate synthesis frame via a feature alignment operation applied to the second and third intermediate synthesis frames (e.g., see the stage 720, FIG. 7). In some examples, generating the fourth and fifth intermediate synthesis frames is performed using a respective plurality of convolution operations.

[0097] In some examples, operations of the block 2108 also include generating a frame for the second set of interpolated frames by fusing the second, fourth and fifth intermediate synthesis frames (e.g., see the stage 730. FIG. 7). In some examples, the fusing of the second, fourth andfifth intermediate synthesis frames is performed using a U-Net merge block (e.g., the block 734, FIG. 7).

[0098] A block 2110 of the method 2100 includes, for a pair of consecutive first frames, generating a third set of interpolated frames at the second frame rate by fusing the first and second sets of interpolated frames generated in the blocks 2106 and 2108, respectively. In some examples, the fusing comprises computing a pairwise weighted sum of corresponding interpolated frames from the first and second sets of interpolated frames. In some examples, a set of weighting coefficients (wl, w2) used for computing the pairwise weighted sum is time dependent.

[0099] In some examples, operations of the block 2110 include (i) generating a first feature map (e.g., 2006, FIG. 20) based on the corresponding interpolated frame (e.g., 612, FIG. 20) from the first set of interpolated frames; and (ii) generating a second feature map (e.g., 2010, FIG. 20) based on the corresponding interpolated frame (e.g., 622, FIG. 20) from the second set of interpolated frames. The set of weighting coefficients (wl, w2) used for computing the pairwise weighted sum is determined based on the first and second feature maps.

[0100] A block 2112 of the method 2100 includes playing, on a playback device, the third set of interpolated frames generated in the block 2110. In some examples, the framerate used in the playback device for playing the third set of interpolated frames is selectable and may be different from the second frame rate.

[0101] FIG. 22 is a block diagram illustrating a computing device 2200 one or more instances of which can be used to implement the method 2100 according to some examples. In some examples, the computing device 2200 is programmed to implement at least some parts of the pipeline 600 (FIGS. 6 and 18). In some examples, the electronic device 100 (FIGS. 1A-1C) may include some or all components of the computing device 2200.

[0102] The computing device 2200 of FIG. 22 is illustrated as having a number of components, but any one or more of these components may be omitted or duplicated, as suitable for the application and setting. In some embodiments, some or all of the components included in the computing device 2200 may be attached to one or more motherboards and enclosed in a housing. In some embodiments, some of those components may be fabricated onto a single system-on-a-chip (SoC) (e.g., the SoC may include one or more electronic processing devices 2202 and one or more storage devices 2204). Additionally, in various embodiments, the computing device 2200 may not include one or more of the components illustrated in FIG. 22,but may include interface circuitry' for coupling to the one or more components using any suitable interface (e.g., a Universal Serial Bus (USB) interface, a High-Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, a wireless interface, or any other appropriate interface). For example, the computing device 2200 may not include a display device 2210, but may include display device interface circuitry (e.g., a connector and driver circuitry ) to which an external display device 2210 may be coupled.

[0103] The computing device 2200 includes a processing device 2202 (e.g., one or more processing devices). As used herein, the terms “electronic processor device” and “processing device" interchangeably refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that may be stored in registers and / or memory7. In various embodiments, the processing device 2202 may include one or more digital signal processors (DSPs), application-specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), server processors, or any other suitable processing devices.

[0104] The computing device 2200 also includes a storage device 2204 (e.g., one or more storage devices). In various embodiments, the storage device 2204 may include one or more memory devices, such as random-access memory (RAM) devices (e.g.. static RAM (SRAM) devices, magnetic RAM (MRAM) devices, dynamic RAM (DRAM) devices, resistive RAM (RRAM) devices, or conductive-bridging RAM (CBRAM) devices), hard drive-based memory devices, solid-state memory7devices, networked drives, cloud drives, or any combination of memory devices. In some embodiments, the storage device 2204 may include memory that shares a die with the processing device 2202. In such an embodiment, the memory may be used as cache memory and include embedded dynamic random-access memory (eDRAM) or spin transfer torque magnetic random-access memory7(STT-MRAM), for example. In some embodiments, the storage device 2204 may include non-transitory computer readable media having instructions thereon that, when executed by one or more processing devices (e.g.. the processing device 2202), cause the computing device 2200 to perform any appropriate ones of the methods disclosed herein below or portions of such methods.

[0105] The computing device 2200 further includes an interface device 2206 (e.g., one or more interface devices 2206). In various embodiments, the interface device 2206 may include one or more communication chips, connectors, and / or other hardware and software to govern communications between the computing device 2200 and other computing devices. Forexample, the interface device 2206 may include circuitry for managing wireless communications for the transfer of data to and from the computing device 2200. The term “wireless” and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data via modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not. Circuitry included in the interface device 2206 for managing wireless communications may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.11 family), IEEE 802.16 standards, Long-Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as “3GPP2”), etc.). In some embodiments, circuitry included in the interface device 2206 for managing wireless communications may operate in accordance with a Global System for Mobile Communication (GSM), General Packet Radio Sendee (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE network. In some embodiments, circuitry included in the interface device 2206 for managing wireless communications may operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). In some embodiments, circuitry7included in the interface device 2206 for managing wireless communications may operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. In some embodiments, the interface device 2206 may include one or more antennas (e.g., one or more antenna arrays) configured to receive and / or transmit wireless signals.

[0106] In some embodiments, the interface device 2206 may include circuitry7for managing wired communications, such as electrical, optical, or any other suitable communication protocols. For example, the interface device 2206 may include circuitry to support communications in accordance with Ethernet technologies. In some embodiments, the interface device 2206 may support both wireless and wired communication, and / or may support multiple wired communication protocols and / or multiple wireless communication protocols. For example, a first set of circuitry of the interface device 2206 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second set of circuitry of theinterface device 2206 may be dedicated to longer-range wireless communications such as global positioning system (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV -DO. or others. In some other embodiments, a first set of circuitry of the interface device 2206 may be dedicated to wireless communications, and a second set of circuitry of the interface device 2206 may be dedicated to wired communications.

[0107] The computing device 2200 also includes battery / power circuitry 2208. In various embodiments, the battery / power circuitry 2208 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 2200 to an energy' source separate from the computing device 2200 (e.g., to AC line power).

[0108] The computing device 2200 also includes a display device 2210 (e.g., one or multiple individual display devices). In various embodiments, the display device 2210 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.

[0109] The computing device 2200 also includes additional input / output (I / O) devices 2212. In various embodiments, the I / O devices 2212 may include one or more data / signal transfer interfaces, audio I / O devices (e.g., microphones or microphone arrays, speakers, headsets, earbuds, alarms, etc.), audio codecs, video codecs, printers, sensors (e.g.. thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration sensors, etc ), image capture devices (e.g., one or more cameras), human interface devices (e.g., keyboards, cursor control devices, such as a mouse, a stylus, a trackball, or a touchpad), etc.

[0110] Depending on the specific embodiment, various components of the interface devices 2206 and / or I / O devices 2212 can be configured to output suitable control signals, receive suitable control / telemetry signals, and receive and transmit data streams. In some examples, the interface devices 2206 and / or I / O devices 2212 include one or more analog-to-digital converters (ADCs) for transforming received analog signals into a digital form suitable for operations performed by the processing device 2202 and / or the storage device 2204. In some additional examples, the interface devices 2206 and / or I / O devices 2212 include one or more digital-to-analog converters (DACs) for transforming digital signals provided by the processing device 2202 and / or the storage device 2204 into an analog form suitable for being transmitted through a communication channel.

[0111] According to an example embodiment disclosed above, e.g.. in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-22, provided is anapparatus comprising: at least one processor; and at least one memory' including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: receive a sequence of first frames representing a scene as captured by a first camera at a first frame rate; receive a sequence of second frames representing the scene as captured by a second camera at a second frame rate that is greater than the first frame rate; and for a pair of consecutive first frames, perform a first interpolation to obtain a first set of interpolated frames at the second frame rate, the first interpolation being based on the pair of consecutive first frames and further based on a corresponding set of the second frames; perform a different second interpolation to obtain a second set of interpolated frames at the second frame rate, the different second interpolation being based on the pair of consecutive first frames and further based on the corresponding set of the second frames; and generate a third set of interpolated frames at the second frame rate by fusing the first and second sets of interpolated frames.

[0112] According to another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-22, provided is a video interpolation method comprising: receiving a sequence of first frames representing a scene as captured by a first camera at a first frame rate; receiving a sequence of second frames representing the scene as captured by a second camera at a second frame rate that is greater than the first frame rate; and for a pair of consecutive first frames, performing a first interpolation to obtain a first set of interpolated frames at the second frame rate, the first interpolation being based on the pair of consecutive first frames and further based on a corresponding set of the second frames; performing a different second interpolation to obtain a second set of interpolated frames at the second frame rate, the different second interpolation being based on the pair of consecutive first frames and further based on the corresponding set of the second frames; and generating a third set of interpolated frames at the second frame rate by fusing the first and second sets of interpolated frames.

[0113] In some embodiments of the above method, the first camera comprises a CMOS pixelated sensor; the second camera comprises a single photon avalanche diode array (SPAD) sensor; and the second frames are binary' frames.

[0114] In some embodiments of any of the above methods, the first and second frame rates differ by a factor of ten or more.

[0115] In some embodiments of any of the above methods, the method further comprises: for each time T of the sequence of second frames, generating a respective stack of virtual exposure frames, each of the virtual exposure frames being an average of N of the second frames nearest to the time T in the sequence, the respective stack including two or more virtual exposure frames corresponding to different respective values of N.

[0116] In some embodiments of any of the above methods, the respective stack includes virtual exposure frames corresponding to N = 1, 2, 4, 8, 16, 32, and 64.

[0117] In some embodiments of any of the above methods, performing the second interpolation for the time T is based on a plurality of the respective stacks corresponding to a set of frames adjacent to the time T.

[0118] In some embodiments of any of the above methods, the method further comprises generating a key frame corresponding to the pair of consecutive first frames by computing a weighted sum of first and second constituent frames of the pair, wherein performing the second interpolation for the time T is further based on the key frame.

[0119] In some embodiments of any of the above methods, performing the second interpolation comprises: for the time T, generating a first intermediate synthesis frame based on the respective stack of virtual exposure frames corresponding to time T-l; generating a second intermediate synthesis frame based on the respective stack of virtual exposure frames corresponding to the time T; and generating a third intermediate synthesis frame based on the respective stack of virtual exposure frames corresponding to time T+l.

[0120] In some embodiments of any of the above methods, each of said generating the first, second, and third intermediate synthesis frames is performed using a respective U-Net synthesis block.

[0121] In some embodiments of any of the above methods, performing the second interpolation further comprises: generating a fourth intermediate synthesis frame via a feature alignment operation applied to the first and second intermediate synthesis frames; and generating a fifth intermediate synthesis frame via a feature alignment operation applied to the second and third intermediate synthesis frames.

[0122] In some embodiments of any of the above methods, each of said generating the fourth and fifth intermediate synthesis frames is performed using a respective plurality of convolution operations.

[0123] In some embodiments of any of the above methods, performing the second interpolation further comprises generating a frame for the second set of interpolated frames by fusing the second, fourth and fifth intermediate synthesis frames.

[0124] In some embodiments of any of the above methods, the fusing of the second, fourth and fifth intermediate synthesis frames is performed using a U-Net merge block.

[0125] In some embodiments of any of the above methods, generating the first, second, and third intermediate synthesis frames comprises deformable convolution (e.g., as expressed by Eqs. (13), (14)).

[0126] In some embodiments of any of the above methods, performing the first interpolation is based on the respective stack corresponding to the time T.

[0127] In some embodiments of any of the above methods, performing the first interpolation comprises warping a key frame corresponding to the pair of consecutive first frames to an intermediate temporal position.

[0128] In some embodiments of any of the above methods, the method further comprises generating the key frame by computing a weighted sum of first and second constituent frames of the pair.

[0129] In some embodiments of any of the above methods, the warping comprises: estimating optical flow between the key frame and an interpolated frame at the intermediate temporal position based on optical flow between a corresponding pair of the second frames; and obtaining the interpolated frame at the intermediate temporal position by warping the key frame using the estimated optical flow.

[0130] In some embodiments of any of the above methods, the estimating is performed using aPWC-Net block.

[0131] In some embodiments of any of the above methods, the estimating and the obtaining are implemented using a neural network trained based on a loss function including an LI loss between a respective training frame at the intermediate temporal position and a correspondingone of the second frames warped towards the training frame using the optical flow (e.g., as expressed by Eq. (22)).

[0132] In some embodiments of any of the above methods, the first interpolation is performed using a neural network comprising a plurality of PWC-Net blocks, a plurality of denoising blocks connected to the plurality of PWC-Net blocks, and a plurality of upsampling layers, with each of the upsampling layers being connected to a respective one of the denoising blocks.

[0133] In some embodiments of any of the above methods, the performing and the generating are implemented using a neural network trained based on a loss function including a weighted sum of an LI loss between a synthesized frame and a corresponding ground truth frame and an LI loss between a spatial gradient of the synthesized frame and a spatial gradient of the corresponding ground truth frame (e.g., as expressed by Eq. (18)).

[0134] In some embodiments of any of the above methods, the fusing is implemented using a neural network comprising an element selected from the group consisting of: a learnable linear layer configured to perform weighted linear averaging, a network configured to perform an average pooling operation or a max pooling operation, a learnable MLP, a learnable CNN, and a learnable transformer.

[0135] In some embodiments of any of the above methods, the fusing comprises computing a pairwise weighted sum of corresponding interpolated frames from the first and second sets of interpolated frames.

[0136] In some embodiments of any of the above methods, a set of weighting coefficients (wl, w2) used for computing the pairwise weighted sum is time dependent.

[0137] In some embodiments of any of the above methods, the computing comprises: generating a first feature map based on the corresponding interpolated frame from the first set of interpolated frames; and generating a second feature map based on the corresponding interpolated frame from the second set of interpolated frames, wherein a set of weighting coefficients (wl, w2) used for computing the pairwise weighted sum is determined based on the first and second feature maps.

[0138] In some embodiments of any of the above methods, the first interpolation includes motion-based frame interpolation; and wherein the second interpolation includes motionindependent frame interpolation.

[0139] In some embodiments of any of the above methods, the method further comprises playing the third set of interpolated frames on a playback device.

[0140] According to yet another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-22, provided is a non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the above methods.

[0141] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other w ords, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.

[0142] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments w ill occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.

[0143] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary' is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.

[0144] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subj ect matter.

[0145] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims.

[0146] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit.

[0147] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes. CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium including being loaded into and / or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits.

[0148] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word ‘‘about” or “approximately” preceded the value or range.

[0149] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be constmed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.

[0150] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence.

[0151] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments necessarily mutually exclusive of other embodiments. The same applies to the term “implementation.”

[0152] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objects so referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner.

[0153] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if’ may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].”

[0154] Also, for purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developedin which energy is allowed to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, although not required. Conversely, the terms '“directly coupled,’7“‘directly connected.” etc., imply the absence of such additional elements.

[0155] As used herein in reference to an element and a standard, the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard and would be recognized by other elements as sufficiently capable of communicating with the other elements in the manner specified by the standard. The compatible element does not need to operate internally in a manner specified by the standard.

[0156] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and / or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and / or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.

[0157] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); (b) combinations of hardw are circuits and softw are, such as (as applicable): (i) a combination of analog and / or digital hardw are circuit(s) w ith software / firmware and (ii) any portions of hardw are processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions): and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e g., firmware) for operation, but the software may not be present when it is not neededfor operation.” This definition of circuitry' applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry' also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

[0158] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may' be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.

Claims

CLAIMSWhat is claimed is:

1. A video interpolation method, comprising:receiving a sequence of first frames representing a scene as captured by a first camera at a first frame rate;receiving a sequence of second frames representing the scene as captured by a second camera at a second frame rate that is greater than the first frame rate; andfor a pair of consecutive first frames.performing a first interpolation to obtain a first set of interpolated frames at the second frame rate, the first interpolation being based on the pair of consecutive first frames and further based on a corresponding set of the second frames;performing a different second interpolation to obtain a second set of interpolated frames at the second frame rate, the different second interpolation being based on the pair of consecutive first frames and further based on the corresponding set of the second frames; andgenerating a third set of interpolated frames at the second frame rate by fusing the first and second sets of interpolated frames.

2. The method of claim 1,wherein the first camera comprises a pixelated CMOS sensor;wherein the second camera comprises a single photon avalanche diode array (SPAD) sensor; andwherein the second frames are binary frames.

3. The method of claim 1 or claim 2, wherein the first and second frame rates differ by a factor of ten or more.

4. The method of any of claims 1-3, further comprising:for each time T of the sequence of second frames, generating a respective stack of virtual exposure frames, each of the virtual exposure frames being an average of N of the second frames nearest to the time T in the sequence, the respective stack including two or more virtual exposure frames corresponding to different respective values of N.

5. The method of claim 4, wherein the respective stack includes virtual exposure frames corresponding to N = 1, 2, 4, 8, 16, 32, and 64.

6. The method of claim 4 or claim 5, wherein performing the second interpolation for the time T is based on a plurality of the respective stacks corresponding to a set of frames adjacent to the time T.

7. The method of any of claims 4-6, further comprising generating a key frame corresponding to the pair of consecutive first frames by computing a weighted sum of first and second constituent frames of the pair,wherein performing the second interpolation for the time T is further based on the key frame.

8. The method of any of claims 4-7, wherein performing the second interpolation comprises:for the time T,generating a first intermediate synthesis frame based on the respective stack of virtual exposure frames corresponding to time T-l;generating a second intermediate synthesis frame based on the respective stack of virtual exposure frames corresponding to the time T; andgenerating a third intermediate synthesis frame based on the respective stack of virtual exposure frames corresponding to time T+l.

9. The method of claim 8. wherein each of said generating the first, second, and third intermediate synthesis frames is performed using a respective U-Net synthesis block.

10. The method of claim 8 or claim 9, wherein performing the second interpolation further comprises:generating a fourth intermediate synthesis frame via a feature alignment operation applied to the first and second intermediate synthesis frames; andgenerating a fifth intermediate synthesis frame via a feature alignment operation applied to the second and third intermediate synthesis frames.

11. The method of claim 10, wherein each of said generating the fourth and fifth intermediate synthesis frames is performed using a respective plurality of convolution operations.

12. The method of claim 10 or claim 11, wherein performing the second interpolation further comprises:generating a frame for the second set of interpolated frames by fusing the second, fourth and fifth intermediate synthesis frames.

13. The method of claim 12, wherein the fusing of the second, fourth and fifth intermediate synthesis frames is performed using a U-Net merge block.

14. The method of any of claims 8-13, wherein generating the first, second, and third intermediate synthesis frames comprises deformable convolution.

15. The method of any of claims 4-14, wherein performing the first interpolation is based on the respective stack corresponding to the time T.

16. The method of any of claims 1-1, wherein performing the first interpolation comprises warping a key frame corresponding to the pair of consecutive first frames to an intermediate temporal position.17 The method of claim 16, further comprising generating the key frame by computing a weighted sum of first and second constituent frames of the pair.

18. The method of claim 16 or claim 17. wherein the warping comprises:estimating optical flow between the key frame and an interpolated frame at the intermediate temporal position based on optical flow between a corresponding pair of the second frames; andobtaining the interpolated frame at the intermediate temporal position by warping the key frame using the estimated optical flow.19 The method of claim 18, wherein the estimating is performed using a PWC-Net block.

20. The method of claim 18, wherein the estimating and the obtaining are implemented using a neural network trained based on a loss function including an LI loss between a respectivetraining frame at the intermediate temporal position and a corresponding one of the second frames warped towards the training frame using the optical flow.

21. The method of any of claims 1-20, wherein the first interpolation is performed using a neural network comprising a plurality of PWC-Net blocks, a plurality of denoising blocks connected to the plurality of PWC-Net blocks, and a plurality of upsampling layers, with each of the upsampling layers being connected to a respective one of the denoising blocks.

22. The method of any of claims 1-21, wherein the performing and the generating are implemented using a neural network trained based on a loss function including a weighted sum of an LI loss between a synthesized frame and a corresponding ground truth frame and an LI loss between a spatial gradient of the synthesized frame and a spatial gradient of the corresponding ground truth frame.

23. The method of any of claims 1-22, wherein the fusing is implemented using a neural network comprising an element selected from the group consisting of: a learnable linear layer configured to perform weighted linear averaging, a network configured to perform an average pooling operation or a max pooling operation, a learnable MLP, a learnable CNN, and a learnable transformer.

24. The method of any of claims 1-23, wherein the fusing comprises computing a pairwise weighted sum of corresponding interpolated frames from the first and second sets of interpolated frames.

25. The method of claim 24, wherein a set of weighting coefficients (wl, w2) used for computing the pairwise weighted sum is time dependent.

26. The method of claim 24, wherein the computing comprises:generating a first feature map based on the corresponding interpolated frame from the first set of interpolated frames; andgenerating a second feature map based on the corresponding interpolated frame from the second set of interpolated frames,wherein a set of weighting coefficients (wl, w2) used for computing the pairwise weighted sum is determined based on the first and second feature maps.

27. The method of any of claims 1-26, wherein the first interpolation includes motion-based frame interpolation; and wherein the second interpolation includes motion-independent frame interpolation.

28. The method of any of claims 1-27, wherein the method further comprises playing the third set of interpolated frames on a playback device.

29. An apparatus comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to perform the method of any of claims 1-28.

30. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform the method of any of claims 1- 28.

Citation Information

Patent Citations

  • High-frame-rate super-resolution improvement method based on multi-modal acquisition

    CN115393194A

  • Association of concurrent tracks across multiple views

    US20230169683A1

  • EP25154438A

  • US202463714445P

  • US202563747674P