Method and apparatus for generating intermediate video frames

The generation of intermediate frames through layer-by-layer recursive call and feature coding sharing solves the problems of low trueness and high complexity of intermediate frames in the existing video interpolation scheme, and realizes efficient and low-complexity intermediate frame generation, which is suitable for high-resolution video interpolation and edge device deployment.

CN115065796BActive Publication Date: 2025-07-25SAMSUNG ELECTRONICS CHINA R&D CENT +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210669349.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-14
Publication Date
2025-07-25
Estimated Expiration
2042-06-14

AI Technical Summary

Technical Problem

In the existing optical flow-based video interpolation scheme, the intermediate frame has low trueness and high implementation complexity, making it difficult to effectively deploy in edge devices.

Method used

Using a layer-by-layer recursive call method, a pre-trained bidirectional optical flow estimation model and pixel synthesis model is used to repair optical flow and intermediate frames in each layer of image, intermediate frames are generated, and a feature encoding network is shared between the optical flow estimation model and pixel synthesis model.

Benefits of technology

It improves the authenticity of the intermediate frame, reduces the number of parameters and implementation complexity of the model, making the solution easier to deploy in edge devices, and is especially suitable for high-resolution video interpolation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115065796B_ABST
    Figure CN115065796B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for generating an intermediate video frame. The method includes: obtaining a target video frame pair, and respectively constructing an image pyramid for each video frame in the target video frame pair; based on the image pyramid, in the order from the highest layer to the lowest layer, by means of a recursive call layer by layer, using a pre-trained bidirectional optical flow estimation model and a pixel synthesis model to generate an intermediate frame of the target video frame pair; wherein, when performing intermediate frame generation processing based on each layer of the image in the image pyramid, based on the current layer of the image, using the optical flow estimation model to repair the bidirectional optical flow obtained from the previous layer of processing, and based on the repaired bidirectional optical flow and the current layer of the image, using the pixel synthesis model to repair the intermediate frame obtained from the previous layer of processing to obtain the intermediate frame output by the current layer of processing. By using the present application, the authenticity of the intermediate frame can be effectively improved and the implementation complexity can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer vision technology, and particularly to a method and apparatus for generating intermediate frames of a video. Background Art

[0002] Video frame interpolation is an important application in computer vision, aiming to improve the frame rate of a video by synthesizing frames that do not exist between consecutive frames (i.e., intermediate frames), thereby making the motion in the video smoother and enhancing the viewing experience of the viewer. For example, for some old videos, limited by the shooting equipment at that time, the frame rate is generally only 25 FPS. Modern high-end TVs usually support a playback speed of 120 FPS. Thus, when playing these low-frame-rate videos on modern TVs, on the one hand, there will be a feeling of lag, and on the other hand, the strongest performance of the TV cannot be exerted. To solve the above problems, artificial intelligence technology can be used to real-time increase the frame rate of the video to 120 FPS by frame interpolation. In addition, video frame interpolation technology also has extensive applications in video compression, view synthesis, adaptive streaming media, etc. In particular, for the applications of the metaverse, the real-time requirement of high-definition video interaction is becoming increasingly important. However, real-time transmission of high-definition and high-frame-rate video streams will bring a very large pressure on network bandwidth. Under limited bandwidth conditions, transmitting a lower-frame-rate video stream and then using video frame interpolation technology on the client side to convert the low-frame-rate video into a high-frame-rate video is a very effective application solution.

[0003] Currently, deep learning algorithms have achieved fruitful results in the field of video frame interpolation. In particular, pixel synthesis based on optical flow is the mainstream implementation method in the current video frame interpolation field. In this type of algorithm, first, the optical flow between the input frame and the target frame is estimated, and then the estimated optical flow is used to guide the synthesis of the intermediate frame. Among them, the optical flow depicts the pixel-level motion between consecutive frames, and its role in the pixel synthesis process is to move the pixels in the input frame to the intermediate frame through forward-warping or backward-warping operations towards the middle. Based on the results of the warping operation, a synthesis network is used to fuse pixel and feature information to generate the final intermediate frame. Generally speaking, the optical flow estimation network is a network with a pyramid structure, which estimates the optical flow from coarse to fine iteratively; the synthesis network is usually a network with an encoder-decoder structure, that is, the commonly mentioned U-Net structure.

[0004] The inventors found in the process of implementing the present application that: the existing video frame interpolation scheme implemented based on optical flow-based pixel synthesis has problems such as low authenticity of intermediate frames and complex implementation. The specific reasons are analyzed as follows:

[0005] In the above video frame interpolation scheme, the intermediate frame is synthesized based on the estimated result of optical flow. Thus, the accuracy of the optical flow estimation result directly determines the accuracy of the generation of the intermediate frame. The existing pyramid structure optical flow model used for estimating optical flow usually only outputs the final result at the highest resolution level, and this result is prone to a large error from the real image. Correspondingly, it will affect the authenticity of the intermediate frame synthesized based on the optical flow estimation result, that is, the error between the intermediate frame and the real image at the corresponding moment is large.

[0006] In addition, the existing video frame interpolation scheme adopts a one-time synthesis method (that is, only running the synthesis network module once), which requires multiple downsamplings during the synthesis process to reduce the inaccuracy of optical flow estimation. In this way, the scale of the synthesis network is large and the number of parameters is very large, which is not conducive to deployment in edge devices in actual application scenarios. Summary of the Invention

[0007] In view of this, the main object of the present invention is to provide a method and device for generating video intermediate frames, which can effectively improve the authenticity of the intermediate frame and reduce the implementation complexity.

[0008] To achieve the above object, the technical solution proposed by the embodiments of the present invention is as follows:

[0009] A method for generating video intermediate frames, including:

[0010] Obtaining a target video frame pair, and respectively constructing an image pyramid for each video frame in the target video frame pair;

[0011] Based on the image pyramid, in the order from the highest layer to the lowest layer, using a recursive call method layer by layer, and using a pre-trained bidirectional optical flow estimation model and a pixel synthesis model to generate the intermediate frame of the target video frame pair; wherein, when performing intermediate frame generation processing based on each layer of the image in the image pyramid, based on the image of the current layer, using the optical flow estimation model to repair the bidirectional optical flow obtained from the processing of the previous layer, and based on the repaired bidirectional optical flow and the image of the current layer, using the pixel synthesis model to repair the intermediate frame obtained from the processing of the previous layer to obtain the intermediate frame output by the processing of the current layer.

[0012] The embodiments of the present invention also propose a device for generating video intermediate frames, including:

[0013] An image pyramid construction module, configured to obtain a target video frame pair and respectively construct an image pyramid for each video frame in the target video frame pair;

[0014] An intermediate frame generation module, configured to generate an intermediate frame of the target video frame pair based on the image pyramid in the order from the highest layer to the lowest layer by means of recursive layer-by-layer calls, using a pre-trained bidirectional optical flow estimation model and a pixel synthesis model; wherein, when performing intermediate frame generation processing based on each layer of image in the image pyramid, based on the current layer of image, the bidirectional optical flow obtained from the previous layer is repaired using the optical flow estimation model, and based on the repaired bidirectional optical flow and the current layer of image, the intermediate frame obtained from the previous layer is repaired using the pixel synthesis model to obtain the intermediate frame output by the current layer of processing.

[0015] An embodiment of the present invention provides a video intermediate frame generation device, including a processor and a memory;

[0016] The memory stores an application program executable by the processor, which is used to cause the processor to execute the above-mentioned video intermediate frame generation method.

[0017] An embodiment of the present invention also provides a computer-readable storage medium, which stores computer-readable instructions for executing the above-mentioned video intermediate frame generation method.

[0018] A computer program product according to an embodiment of the present invention includes a computer program / instructions, characterized in that when the computer program / instructions are executed by a processor, the steps of the above-mentioned video intermediate frame generation method are implemented.

[0019] In summary, in the video intermediate frame generation solution proposed by the embodiment of the present invention, when performing intermediate frame generation processing based on each layer of image in the image pyramid, it is necessary to repair the optical flow obtained from the previous layer of processing, and based on the bidirectional optical flow repaired in the current layer and the image of the image pyramid in the current layer, repair the intermediate frame obtained from the previous layer of intermediate frame generation processing. In this way, by repairing the optical flow and the intermediate frame during the generation processing of the intermediate frame based on each layer of image, rather than only using the optical flow obtained during the processing of the highest resolution level to repair the intermediate frame, the robustness of the optical flow estimation quality is effectively improved, thereby effectively improving the accuracy of the intermediate frame, that is, the generated intermediate frame can be closer to the real image at the corresponding moment, and further effectively improving the interpolation visual effect, especially significantly improving the interpolation effect for high-resolution videos.

[0020] In addition, by repairing the intermediate frames in each layer of processing, the synthesis of intermediate frames needs to be executed multiple times. In this way, the number of samplings required for each synthesis can be effectively reduced, thereby reducing the scale of the pixel synthesis model and significantly reducing the number of model parameters. Moreover, by adopting a recursive call method layer by layer to generate intermediate frames, each layer of processing can share the model. In this way, the network parameters of the entire solution are also effectively reduced. Therefore, adopting the embodiments of the present invention can effectively reduce the implementation complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic flowchart of the method according to the embodiment of the present invention;

[0022] Figure 2 is a schematic diagram of the bidirectional optical flow estimation model according to the embodiment of the present invention;

[0023] Figure 3 is a schematic diagram of the pixel synthesis model according to the embodiment of the present invention;

[0024] Figure 4 is a schematic diagram for comparing the optical flow estimation effects between the embodiment of the present invention and the existing PWC-Net algorithm;

[0025] Figure 5 is a schematic diagram of the structure of the video intermediate frame generation model according to the embodiment of the present invention;

[0026] Figure 6 is a schematic diagram for comparing the single-frame video interpolation examples and the comparison between the embodiment of the present invention and the existing solution;

[0027] Figure 7 is a schematic diagram of the multi-frame video interpolation example according to the embodiment of the present invention;

[0028] Figure 8 is a schematic diagram of the device structure according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] Figure 1 is a schematic flowchart of the method according to the embodiment of the present invention. As Figure 1 shown, the video intermediate frame generation method implemented in this embodiment mainly includes:

[0031] Step 101: Obtain a pair of target video frames, and respectively construct an image pyramid for each video frame in the pair of target video frames.

[0032] In this step, to facilitate improving the efficiency of subsequent optical flow estimation, for each video frame in the target video frame pair (specifically composed of two consecutive video frames) of the intermediate frame to be generated currently, an image pyramid is respectively constructed, so that in the subsequent steps, based on each layer of the image in the image pyramid, the generation process of the intermediate frame is performed layer by layer.

[0033] The construction of the image pyramid in this step can be implemented by using existing methods. The specific number of layers of the image pyramid is related to the resolution of the video frames in the target video frame pair. The larger the resolution, the more layers the pyramid has.

[0034] In practical applications, the pyramid level in the test stage can be increased to facilitate predicting an ultra-large optical flow beyond the training set. For example, when training the model of the embodiment of the present invention on the low-resolution Vimeo90K, only three levels of the pyramid are sufficient to predict the optical flow on the training set. However, if predictions are made on an ultra-high-resolution (4K) dataset such as 4K1000FPS, it is necessary to increase the pyramid level during testing. For example, it is recommended to use seven levels.

[0035] Step 102: Based on the image pyramid, in the order from the highest layer to the lowest layer, by using a recursive call layer by layer, an intermediate frame of the target video frame pair is generated by using a pre-trained bidirectional optical flow estimation model and a pixel synthesis model; wherein, when performing the generation process of the intermediate frame based on each layer of the image in the image pyramid, based on the image of the current layer, the bidirectional optical flow obtained by the processing of the previous layer is repaired by using the optical flow estimation model, and based on the repaired bidirectional optical flow and the image of the current layer, the intermediate frame obtained by the processing of the previous layer is repaired by using the pixel synthesis model to obtain the intermediate frame output by the processing of the current layer.

[0036] This step is used to generate an intermediate frame between the two video frames by using a pre-trained bidirectional optical flow estimation model and a pixel synthesis model based on the image pyramids of the two video frames obtained in step 101. Here, the moment corresponding to the intermediate frame is between the two video frames and is not limited to the exact middle moment.

[0037] It should be noted that when generating intermediate frames based on each layer of images, the optical flow and intermediate frames obtained from the previous level need to be repaired, rather than only using the optical flow obtained during the processing at the highest resolution level to repair the intermediate frames. In this way, the robustness of the optical flow estimation quality can be effectively improved, thereby effectively improving the authenticity of the intermediate frames, and further effectively improving the frame interpolation effect, especially significantly improving the frame interpolation effect for high-resolution videos. Moreover, by repairing the intermediate frames in each layer of processing, the scale of the pixel synthesis model can be greatly reduced, and the number of model parameters can be reduced. In addition, by adopting a recursive call method layer by layer, the network parameter scale of the entire intermediate frame generation scheme can be effectively reduced, thereby effectively reducing the implementation complexity and improving the generation efficiency of the intermediate frames.

[0038] In step 102, the intermediate frame obtained by performing intermediate frame generation processing based on the image of the last layer of the pyramid is the intermediate frame of the target video frame pair finally obtained.

[0039] Furthermore, the inventor found during the implementation of the present invention that: in the existing video frame interpolation schemes, both the optical flow network and the synthesis network use a feature encoding module for the input pictures in their designs. That is, when the optical flow estimation model performs optical flow estimation, it needs to construct a matching cost volume based on pixel-level features, and the pixel synthesis model requires pixel-level feature maps to provide context information. Therefore, the feature encoding network for generating feature maps for images can be shared by the bidirectional optical flow estimation model and the pixel synthesis model. However, in the existing schemes, the optical flow network and the synthesis network operate independently. Correspondingly, both independently use their own feature encoding modules, which leads to redundant configuration of the modules, and further leads to a large parameter scale of the scheme, increasing the implementation complexity of the scheme. To address the above problems, in order to further reduce the parameter scale and the implementation complexity, in one implementation manner, the bidirectional optical flow estimation model and the pixel synthesis model can share the feature encoding network to reduce the model redundancy, and at the same time, the robustness of generating feature maps for each layer of images can also be improved. Specifically, it can be implemented by the following method:

[0040] When performing intermediate frame generation processing based on each layer of images in the image pyramid, before using the optical flow estimation model to repair the bidirectional optical flow obtained from the previous layer of processing, use a preset feature encoding network to generate a first number of pixel-level feature maps with different resolutions for each image in the current layer of the image pyramid, respectively, to be provided to the optical flow estimation model and the pixel synthesis model for the respective repairs.

[0041] Among them, the first quantity is greater than or equal to 3; the feature encoding network is a convolutional network with at least a second quantity of downsamplings; the second quantity is equal to the first quantity minus one. Specifically, those skilled in the art can set reasonable values of the first quantity and the second quantity according to actual needs. For example, when the first quantity is 3, the second quantity is 2, that is, when the feature encoding network has two downsamplings, it will output pixel-level feature maps with three different resolutions.

[0042] In one implementation, the feature encoding network can be specifically implemented by a 12-layer convolutional network, which is a convolutional network with at least two downsamplings and outputs feature maps with three different resolutions. This feature encoding network performs downsamplings at the 5th layer and the 9th layer respectively, and finally outputs feature maps with three different resolutions (that is, the feature maps before each downsampling and the feature map output by the last layer of convolution), that is, the feature maps from the 4th layer, the 8th layer, and the 12th layer of the convolutional network. The above is only an exemplary description of the specific implementation of the feature encoding network, and it is not limited to this in actual applications.

[0043] In one implementation, in step 102, when performing intermediate frame generation processing based on each layer of the image in the image pyramid, the following method can be specifically used to repair the bidirectional optical flow obtained from the previous layer of intermediate frame generation processing:

[0044] When performing intermediate frame generation processing based on each layer of the image in the image pyramid, the pixel-level feature map corresponding to the current layer of the image and the bidirectional optical flow obtained from the previous layer of processing are input into the bidirectional optical flow estimation model for optical flow repair processing. Among them, the pixel-level feature map is the feature map output by the last layer of convolution of the feature encoding network when the current layer of the image is input into the preset feature encoding network for processing; the bidirectional optical flow is the optical flow from each video frame in the target video frame pair to the current intermediate frame to be generated.

[0045] In one implementation, as Figure 2 shown, the bidirectional optical flow estimation model in the above method can specifically perform optical flow repair processing using the following steps:

[0046] Step a1: Linearly weight the bidirectional optical flow obtained from the previous layer of processing to obtain an initial estimate of the bidirectional optical flow for the current layer of processing.

[0047] Here, the parameter regarding time used for linear weighting is the time corresponding to the current intermediate frame to be generated. The specific method of linear weighting is mastered by those skilled in the art and will not be elaborated here.

[0048] In particular, when performing the first layer of intermediate frame generation processing, the object of linear weighting, that is, the initial value of the bidirectional optical flow, can be set to 0.

[0049] Step a2: Based on the initial bidirectional optical flow estimate, use the forward warping layer of the bidirectional optical flow estimation model to perform forward warping towards the middle on the pixel-level feature maps corresponding to each of the images in the current layer respectively.

[0050] It should be noted here that considering that the existing forward-warping method for multi-frame interpolation (generating multiple frames between two input images) has good support and the overall algorithm framework is more concise, the forward-warping method is selected in the embodiments of the present invention to perform warping processing to align the pixels in the two feature maps corresponding to the two video frames in the target video frame pair.

[0051] Specifically, in step a2, the pixel-level feature maps corresponding to each of the images in the current layer are the feature maps output by the last convolutional kernel of the feature encoding network.

[0052] Step a3: Based on the feature maps obtained by the forward warping, use the matching cost layer of the bidirectional optical flow estimation model to construct a local matching cost.

[0053] In this step, it is used to construct a local matching cost (partial cost volume) by using the cost volume layer. Here, it should be noted that the cost volume is used to represent the matching scores of the pixel-level features of two input images, and the partial cost volume is a very discriminative representation for the optical flow task, specifically referring to the matching scores of the features of the local pixel blocks corresponding to the pixels of one image and another image.

[0054] Here, the specific implementation of constructing the local matching cost is mastered by those skilled in the art and will not be elaborated here.

[0055] Step a4: Based on the initial bidirectional optical flow estimate, the feature maps obtained by the forward warping, the local matching cost, and the CNN features of the bidirectional optical flow obtained during the processing of the previous layer, perform channel stacking; input the result of the channel stacking into the optical flow estimation layer of the bidirectional optical flow estimation model to perform optical flow estimation, and obtain the bidirectional optical flow repair result of the current layer.

[0056] Specifically, the specific implementation of the optical flow estimation is mastered by those skilled in the art, and the optical flow estimation layer can be implemented by a 6-layer convolutional network, but is not limited thereto.

[0057] In one implementation, such as Figure 3As shown, in step 102, when generating an intermediate frame based on each layer of images in the image pyramid, the following method can be specifically used to repair the intermediate frame obtained from the previous layer using a pixel synthesis model:

[0058] Step b1: Linearly weight the bidirectional optical flow obtained from the repair.

[0059] Step b2: For each video frame, use the forward warping layer of the pixel synthesis model to perform forward warping towards the middle on the image at the current layer and the context features of the image based on the linearly weighted optical flow corresponding to the video frame.

[0060] Among them, the context features are the feature maps before each downsampling and the feature map output by the last convolutional layer of the feature encoding network after inputting the image at the current layer of the video frame into the feature encoding network for processing.

[0061] Step b3: Input the result of the forward warping and the intermediate frame obtained from the previous layer into the pixel synthesis network of the pixel synthesis model for processing to obtain the intermediate frame repair result at the current layer.

[0062] In an implementation manner, the pixel synthesis network can be specifically implemented using a simple U-Net structure with only two downsampling layers (composed of an encoding layer and a decoding layer). Specifically, the input of the first encoding layer of the U-Net includes the intermediate frame obtained from the previous layer, the input picture after warping processing, and the context features after warping processing (i.e., the feature maps output by the feature encoding network). The input of the next layer after each downsampling of the encoding layer will embed the context features after warping processing corresponding to the resolution. The U-Net includes a total of two downsamplings and corresponding two upsamplings. Through the decoding network including two upsamplings, the repaired intermediate frame of this pyramid layer can be obtained.

[0063] In practical applications, particularly for the first layer of processing, the initial value of the intermediate frame can be the averaged result after warping of the two images at the current layer.

[0064] Furthermore, considering the above method embodiments, in each layer of processing, the optical flow obtained in the previous layer is repaired, so that a very accurate optical flow estimation result can be finally obtained. In practical applications, the optical flow of the intermediate frame can be applied not only to video frame interpolation, but also to moving object detection, video salient region detection, etc. In particular, since the above embodiments use a flexible pyramid recursive structure, it can flexibly process small optical flow, large optical flow, and complex non-linear optical flow. In this way, a more accurate optical flow estimation value can be obtained by using the above method embodiments, that is, after obtaining the intermediate frame based on the lowest layer image in the image pyramid, a bidirectional optical flow obtained by repairing based on the image of the current layer is output.

[0065] As Figure 4 shown, compared with the commonly used PWC-Net algorithm in the prior art, the method embodiment of the present invention can obtain more accurate and efficient results when estimating complex large optical flow, and has obvious application advantages.

[0066] Figure 5 It is a schematic diagram of a preferred structure of the video intermediate frame generation model corresponding to the above method embodiments. As Figure 5 shown, the video intermediate frame generation model is a pyramid recursive network. This structure enables the present application to use a specified pyramid level in the test stage and flexibly process input pictures of different resolutions. The same recursive unit is repeatedly executed at different levels of the image pyramid, and the finally estimated intermediate frame is output at the highest resolution level. To construct the recursive unit, a feature encoding network is first used to extract pixel-level features of the input picture, and then a bidirectional optical flow estimation model and a pixel synthesis model are used to correct the optical flow and intermediate frame estimated in the previous level respectively. In particular, forward-warping is used to compensate for the optical flow between frames. Since the present application embodiment estimates the bidirectional optical flow between input pictures, the optical flow from the input frame to the intermediate frame required for forward-warping can be easily obtained by means of linear weighting. Therefore, the method of the present application embodiment can easily estimate a certain frame at any time between two input video frames.

[0067] As can be seen from the above method embodiments, in the video intermediate frame generation scheme proposed in this application, when generating intermediate frames based on each layer of images, optical flow and intermediate frame repair are performed, which can effectively improve the robustness of optical flow estimation quality, thereby effectively improving the authenticity of intermediate frames, and further effectively improving the frame interpolation effect. Moreover, by performing intermediate frame repair in each layer of processing, the scale of the pixel synthesis model can be greatly reduced, and the number of model parameters can be reduced. In addition, by adopting a model that is recursively called layer by layer to generate intermediate frames, the network parameter scale of the entire implementation scheme is also effectively reduced, thereby effectively reducing the implementation complexity and improving the generation efficiency of intermediate frames. Further, by sharing the feature encoding network between the optical flow estimation model and the pixel synthesis model, the model architecture of the above method embodiments is very lightweight, and the number of parameters is less than 1 / 10 of the existing mainstream solutions.

[0068] Based on the above technical advantages, under the application requirements with high real-time requirements or low computational power consumption, the above method embodiments have strong application potential. Specifically, based on the very lightweight advantages of the above method embodiments, the above method embodiments can be applied to video frame interpolation schemes on terminals. In addition, the above method embodiments can also be applied to downstream tasks based on video frame interpolation, such as video compression, view synthesis, etc. The above method embodiments can also be applied to relieve the transmission pressure of video data in metaverse projects, that is, transmit low-frame-rate videos and restore high-frame-rate videos through frame interpolation technology. In addition, the advantages of the above method embodiments in improving the accuracy of bidirectional optical flow estimation values can also be utilized and applied to various optical flow-related applications.

[0069] Next, the technical effects of the above method embodiments will be further elaborated in combination with specific implementation examples of single-frame video frame interpolation and multi-frame video frame interpolation.

[0070] I. Single-frame video frame interpolation

[0071] A common requirement for video frame interpolation is to synthesize an intermediate frame between two consecutive video frames (frame-0 and frame-1), that is, frame-0.5. As Figure 6As shown in the figure, starting from the left, the first column is the input frame with two video frames overlapped, the second column is the target intermediate frame (i.e., the true intermediate frame, ground truth), the third and fourth columns are the intermediate frames obtained by using two current state-of-the-art video frame interpolation algorithms (AdaCoF and ABME) respectively, and the fifth column is the intermediate frame obtained by using the method embodiment of the present invention. The second row shows the effect after magnifying the local picture. Based on this row of images, it is easier to see the subtle differences between the intermediate frames synthesized by different algorithms. As can be seen from the figure, in the case of local complex non-linear optical flow (the second row), good results can be obtained by using the method embodiment of the present invention.

[0072] II. Multi-frame video frame interpolation

[0073] If the frame rate of the original video is too low (such as 25 FPS), it is often necessary to insert multiple frames between two consecutive frames to make it reach a higher frame rate (such as 120 FPS). Based on the linear weighting method, the optical flow from the input frame to the intermediate frame at any time can be approximately obtained, and then combined with the forward-warping method, the intermediate frame at any time can be synthesized. When multiple intermediate frames need to be generated, the steps in the above method embodiment are executed the corresponding number of times, and the corresponding number of intermediate frames can be obtained. Figure 7 The figure shows an example diagram of multi-frame video frame interpolation. The first row is the target intermediate frame, and the second row is the intermediate frames at different times obtained by using the method embodiment of the present invention. As can be seen from the figure, the multiple intermediate frames obtained by using the method embodiment of the present invention are very close to the target intermediate frame.

[0074] Based on the above method embodiment, the embodiment of the present invention proposes a device for generating video intermediate frames, as Figure 8 shown, including:

[0075] An image pyramid construction module 801, configured to obtain a target video frame pair and respectively construct an image pyramid for each video frame in the target video frame pair.

[0076] An intermediate frame generation module 802, configured to generate an intermediate frame of the target video frame pair based on the image pyramid in the order from high to low of the pyramid layers by using a recursive call layer by layer, and using a pre-trained bidirectional optical flow estimation model and a pixel synthesis model; wherein, when performing intermediate frame generation processing based on each layer of the image in the image pyramid, based on the current layer of the image, the bidirectional optical flow obtained by the processing of the previous layer is repaired by using the optical flow estimation model, and based on the repaired bidirectional optical flow and the current layer of the image, the intermediate frame obtained by the processing of the previous layer is repaired by using the pixel synthesis model to obtain the intermediate frame output by the processing of the current layer.

[0077] It should be noted that the above methods and apparatuses are based on the same inventive concept. Since the principles for the methods and apparatuses to solve problems are similar, the implementation of the apparatus and the method can be referred to each other, and the repeated parts will not be elaborated.

[0078] Based on the above method embodiments, embodiments of the present invention further propose a device for generating an intermediate video frame, including a processor and a memory; an application program executable by the processor is stored in the memory, and is used to enable the processor to execute the method for generating an intermediate video frame as described above. Specifically, a system or device with a storage medium can be provided, and software program codes for implementing the functions of any one of the above embodiments are stored on the storage medium, and the computer (or CPU or MPU) of the system or device reads and executes the program codes stored in the storage medium. In addition, some or all of the actual operations can also be completed by an operating system operating on the computer based on instructions of the program codes. The program codes read from the storage medium can also be written into a memory provided in an expansion board inserted into the computer or into a memory provided in an expansion unit connected to the computer, and then the CPU etc. installed on the expansion board or the expansion unit execute some and all of the actual operations based on the instructions of the program codes, so as to implement the functions of any one of the embodiments of the method for generating an intermediate video frame as described above.

[0079] Among them, the memory can be specifically implemented as various storage media such as an electrically erasable programmable read-only memory (EEPROM), a flash memory, and a programmable read-only memory (PROM). The processor can be implemented as including one or more central processing units or one or more field programmable gate arrays, and one or more central processing unit cores are integrated in the field programmable gate array. Specifically, the central processing unit or the central processing unit core can be implemented as a CPU or an MCU.

[0080] Embodiments of the present application implement a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method for generating an intermediate video frame as described above are implemented.

[0081] It should be noted that not all steps and modules in the above-mentioned processes and structure diagrams are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted according to needs. The division of each module is only for the convenience of description and is a functional division. In actual implementation, one module can be implemented by multiple modules, and the functions of multiple modules can also be implemented by the same module. These modules can be located in the same device or in different devices.

[0082] The hardware modules in each embodiment can be implemented mechanically or electronically. For example, a hardware module may include a specially designed permanent circuit or logic device (such as a dedicated processor, such as an FPGA or ASIC) for performing a specific operation. The hardware module may also include a programmable logic device or circuit (such as a general-purpose processor or other programmable processor) temporarily configured by software to perform a specific operation. As for whether to implement the hardware module mechanically, or using a dedicated permanent circuit, or using a temporarily configured circuit (such as configured by software), it can be decided based on cost and time considerations.

[0083] In this article, "schematic" means "serving as an example, instance or explanation", and any diagram or implementation method described as "schematic" in this article should not be interpreted as a more preferred or more advantageous technical solution. In order to make the drawings concise, only the parts related to the present invention are schematically shown in each figure, and do not represent the actual structure of the product. In addition, in order to make the drawings concise and easy to understand, in some figures, only one of the parts with the same structure or function is schematically drawn, or only one of them is marked. In this article, "one" does not mean that the number of the relevant parts of the present invention is limited to "only one", and "one" does not mean that the number of the relevant parts of the present invention is "more than one". In this article, "upper", "lower", "front", "back", "left", "right", "inside", "outside", etc. are only used to indicate the relative position relationship between the relevant parts, rather than to limit the absolute position of these relevant parts.

[0084] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for generating an intermediate video frame, characterized in that, Including: Obtain a target video frame pair, and respectively construct an image pyramid for each video frame in the target video frame pair; Based on the image pyramid, in the order from the highest layer to the lowest layer, by means of layer-by-layer recursive call, use a pre-trained bidirectional optical flow estimation model and a pixel synthesis model to generate an intermediate frame of the target video frame pair; wherein, when performing intermediate frame generation processing based on each layer of image in the image pyramid, based on the image of the current layer, use the optical flow estimation model to repair the bidirectional optical flow obtained in the previous layer, and based on the repaired bidirectional optical flow and the image of the current layer, use the pixel synthesis model to repair the intermediate frame obtained in the previous layer to obtain the intermediate frame output by the processing of the current layer; Wherein, the method further includes: When performing intermediate frame generation processing based on each layer of image in the image pyramid, before using the optical flow estimation model to repair the bidirectional optical flow obtained in the previous layer, use a preset feature encoding network to respectively generate a first number of pixel-level feature maps with different resolutions for each image in the current layer of the image pyramid to be provided to the optical flow estimation model and the pixel synthesis model for the respective repairs; wherein, the first number is greater than or equal to 3; the feature encoding network is a convolutional network with at least a second number of downsamplings; the second number is equal to the first number minus one.

2. The method according to claim 1, wherein Repairing the bidirectional optical flow obtained in the previous layer includes: When performing intermediate frame generation processing based on each layer of image in the image pyramid, input the pixel-level feature map corresponding to the image of the current layer and the bidirectional optical flow obtained in the previous layer into the bidirectional optical flow estimation model for optical flow repair processing; wherein, the pixel-level feature map is the feature map output by the last layer of convolution of the feature encoding network when the image of the current layer is input to the preset feature encoding network for processing; the bidirectional optical flow is the optical flow from each video frame to the intermediate frame.

3. The method according to claim 2, wherein, The optical flow repair processing includes: Linearly weight the bidirectional optical flow obtained in the previous layer to obtain an initial estimate value of the bidirectional optical flow for the current layer processing; Based on the initial estimate value of the bidirectional optical flow, use the forward warping layer of the bidirectional optical flow estimation model to respectively perform forward warping towards the middle on the pixel-level feature map corresponding to each image of the current layer; Based on the feature map obtained by the forward warping, use the matching cost layer of the bidirectional optical flow estimation model to construct a local matching cost; Based on the initial estimate value of the bidirectional optical flow, the feature map obtained by the forward warping, the local matching cost, and the convolutional neural network (CNN) feature of the bidirectional optical flow obtained in the previous layer processing, perform channel stacking; input the result of the channel stacking into the optical flow estimation layer of the bidirectional optical flow estimation model for optical flow estimation to obtain the bidirectional optical flow repair result of the current layer.

4. The method according to claim 1, characterized in that, Repairing the intermediate frame obtained in the previous layer includes: Linearly weight the repaired bidirectional optical flow; For each of the video frames, the forward warping layer of the pixel synthesis model is used to perform a forward warping toward the middle on the image of the video frame at the current layer and the context features of the image based on the linear weighted optical flow corresponding to the video frame; the context features are feature maps before each downsampling output by the feature encoding network and feature maps output by the last convolution layer after the image of the video frame at the current layer is input into a preset feature encoding network for processing; The result of the forward warping and the intermediate frame obtained by the processing of the previous layer are input into the pixel synthesis network of the pixel synthesis model for processing to obtain the intermediate frame repair result at the current layer.

5. The method according to claim 1, wherein The method further comprises: after obtaining the intermediate frame based on the lowest layer image in the image pyramid, outputting the bidirectional optical flow obtained by performing the restoration based on the image of the current layer.

6. A device for generating an intermediate video frame, characterized in that, include: An image pyramid construction module is used to obtain a target video frame pair and construct an image pyramid for each video frame in the target video frame pair; An intermediate frame generation module is used to generate an intermediate frame of the target video frame pair based on the image pyramid, in a descending order of the pyramid layers, by recursive calling layer by layer, using a pre-trained bidirectional optical flow estimation model and a pixel synthesis model; wherein, when performing intermediate frame generation processing based on each layer of the image in the image pyramid, based on the image of the current layer, using the optical flow estimation model, the bidirectional optical flow obtained by the processing of the previous layer is repaired, and based on the repaired bidirectional optical flow and the image of the current layer, using the pixel synthesis model, the intermediate frame obtained by the processing of the previous layer is repaired to obtain the intermediate frame of the processing output of the current layer; Among them, the intermediate frame generation module is further used to generate a first number of pixel-level feature maps of different resolutions for each image in the current layer of the image pyramid, using a preset feature encoding network, when performing intermediate frame generation processing based on each layer of images in the image pyramid, before using the optical flow estimation model to repair the bidirectional optical flow obtained by the processing of the previous layer, so as to provide the optical flow estimation model and the pixel synthesis model for the repair respectively; wherein the first number is greater than or equal to 3; the feature encoding network is a convolutional network with at least a second number of downsampling times; the second number is equal to the first number minus one.

7. A device for generating intermediate video frames, characterized in that, including a processor and a memory; The memory stores an application program executable by the processor, which is used to enable the processor to execute the method for generating an intermediate video frame as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored therein, and the computer-readable instructions are used to execute the method for generating a video intermediate frame as described in any one of claims 1 to 5.

9. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by a processor, the steps of the method for generating an intermediate frame of a video as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Bidirectional optical flow estimation method and device

    CN114581493A

  • System, device and method for video frame interpolation using structured neural network

    WO2021093432A1