A Video Frame Interpolation Method and System at Any Time Based on Bidirectional Meta-Learning
Through the video interpolation method based on bidirectional meta-learning, arbitrary interpolation is achieved using similar video frame training models, which solves the problem of inability to interpolate at any time and poor adaptability in traditional methods, and improves the video viewing experience.
Patent Information
- Application Number
- CN202210437295.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-19
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-04-19
AI Technical Summary
The existing video interpolation method cannot perform interpolation at any time, and the fixed parameters method has poor adaptability to diversified video content, and the reliability of optical flow is difficult to ensure.
The video interpolation method based on bidirectional meta-learning is adopted to train the video interpolation network model using similar but not adjacent video frames, including an optical flow estimator and a video interpolation generator based on meta-learning, and the interpolation frame is realized at any time through optical flow generation and frame generation meta-learning.
The video frame insertion at any time is realized, which improves the video frame rate and improves the visual experience of the video viewer, and the effect is better than traditional methods.
Smart Images

Figure CN115037902B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of video effect enhancement, and particularly relates to a method and system for video frame interpolation at any time based on bidirectional meta-learning. Background Art
[0002] Video effect enhancement aims to post-process videos with less-than-ideal quality through technical means to improve the viewing experience of video viewers. Video frame interpolation is an important topic in video effect enhancement. In recent years, video frame interpolation has attracted increasing attention in the academic community.
[0003] Traditional video frame interpolation methods can be divided into three categories: 1) direct generation methods based on context frames, which directly attempt to generate the frame to be generated at the moment to be generated by using the frames before and after that moment and a given algorithm; 2) optical flow-guided methods, which use optical flow to align the frames before and after and attempt to generate the frame to be generated; 3) methods based on adaptive convolution kernels, which regard the frame to be generated as the convolution result of adjacent frames, obtain the convolution kernel by learning the frames of the original video, and then generate the frame to be generated.
[0004] However, the above three methods have some common defects. First, whether to use optical flow is a dilemma because reliable optical flow can assist us in generating reliable interpolated frames, but the reliability of the optical flow itself is difficult to guarantee. Second, since the trained parameters are fixed, all of the above methods can only perform frame interpolation at fixed times. Finally, all modules use fixed parameters, and such parameters have poor adaptability to diverse video content. Summary of the Invention
[0005] Aiming at the above technical problems, the purpose of the present invention is to provide a method and system for video frame interpolation at any time based on bidirectional meta-learning. The present invention can perform frame interpolation at any time by means of the frames of a given video, realize the improvement of the video frame rate, and improve the visual experience of video viewers.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A video frame interpolation method based on bidirectional meta-learning. The method includes the following steps:
[0008] Using the existing video as the training data set, train the video frame interpolation network model proposed by the present invention. The model takes two adjacent but non - adjacent video frames at any two close times as input, outputs the video frame at one of these two times, and uses the real video frame at this time as the supervision. The video frame interpolation network model includes two major parts: an optical flow estimator based on meta - learning and a video frame interpolation generator. The optical flow estimator based on meta - learning includes the following parts: a preliminary optical flow estimator, a linear optical flow mixer, an optical flow kernel generator, a feature extractor, a weight estimator, and a meta - optical flow refiner; the video frame interpolation generator based on meta - learning includes an optical - flow - based frame preliminary generator, a frame kernel generator, a feature extractor, and a meta - frame regressor.
[0009] The training method of this model is described as follows.
[0010] Input two adjacent but non - adjacent frames in the training video and the time of the frame to be interpolated into the neural network model to be trained, and perform the following steps from the perspective of optical flow generation and frame interpolation respectively:
[0011] From the perspective of optical flow generation, use optical flow estimation based on meta - learning. The specific steps are as follows:
[0012] Input the images I t 、I t+1 of the aforementioned two non - adjacent frames into the preliminary optical flow estimator, and respectively preliminarily estimate the optical flow from the previous frame to the next frame and the optical flow from the next frame to the previous frame; where I t is before I t+1 .
[0013] Downsample the images of the aforementioned two frames to obtain and input them into the preliminary optical flow estimator, and respectively estimate the optical flow from the previous frame to the next frame and the optical flow from the next frame to the previous frame in the case of downsampling; then upsample the preliminary estimation result back to the original resolution, denoted as
[0014] For the preliminary estimation result of the optical flow at the original resolution and the result of upsampling the preliminary estimation of the optical flow under downsampling back to the original resolution, respectively use the method of linear optical flow mixing to directly calculate the optical flow from any time α of the frame to be interpolated between the two frames to the known frame times, denoted as This α is a decimal between 0 and 1, indicating the relative position of the time of the frame to be interpolated between two known times. For example, α = 0.75 means that the time of the frame to be interpolated is at the upper quartile of the two known times.
[0015] Input the obtained in the previous step jointly with the specific time value t + α of the frame to be interpolated into the weight estimator, and then input the result output by the weight estimator into the optical flow kernel generator to generate the optical flow kernel The specific structures of the weight estimator and the optical flow kernel generator are described in detail later.
[0016] Meanwhile, the preliminary estimated
[0017] is input into the feature extractor to obtain the abstract features of the optical flow. That is the refined optical flow regularization term L is calculated for the preliminary estimated optical flow and the refined optical flow f , and the smoothing loss term L is calculated for the refined optical flow s . The specific calculation methods of the above two items are described in detail later.
[0018] From the perspective of video frame interpolation, a video frame interpolation generator based on meta-learning is used. The specific steps are as follows:
[0019] The above two frames of images I t , I t+1 and the input frame preliminary generator obtained previously are used to obtain the frame at any moment α between the two preliminarily generated frames to be interpolated. Denote the frame preliminary generator as W, that is The reconstruction frame consistency regularization term L is calculated for the generated frame at any moment. w . The specific calculation method of this term and the specific structure of the frame preliminary generator are described in detail later.
[0020] The refined optical flow is jointly input into the weight estimator together with the specific moment value of the frame to be interpolated, and then the result output by the weight estimator is input into the frame kernel generator to generate the frame kernel. The specific structures of the weight estimator and the frame kernel generator are described in detail later.
[0021] Meanwhile, the preliminarily generated frame is jointly used with the two known frames I t , I t+1 to generate the abstract features of the frame using the feature extractor.
[0022] Finally, in the meta-frame regressor, the abstract features of the frame are convolved with the frame kernel to obtain the refined frame. That is This frame is the interpolation result at any moment desired by the present invention. Using the real frame at this moment of the video as the supervision, the reconstruction loss function term L is calculated. r This term L r 's specific calculation method is described in detail later.
[0023] Furthermore, both the training dataset and the cross-test dataset include publicly available video datasets. During training, the training dataset is input into the system for backpropagation and gradient descent, and the effectiveness is verified on the cross-test dataset.
[0024] Furthermore, the aforementioned weight estimators are each composed of 2 convolutional layers and 2 fully connected layers.
[0025] Furthermore, each of the aforementioned feature extractors is a residual dense network. Each network includes 12 residual dense blocks. Each residual dense block contains 8 convolutional layers in sequence, and each convolutional layer contains 64 channels.
[0026] Furthermore, each of the aforementioned kernel generators is also a residual dense network. Each network includes 4 residual dense blocks. Each residual dense block contains 4 convolutional layers in sequence, and each convolutional layer contains 32 channels.
[0027] Furthermore, the processing networks for optical flow features and optical flow kernels, as well as the processing networks for frame features and frame kernels, each include 2 convolutional layers and 2 fully connected layers.
[0028] Combining the aforementioned four losses with weights and performing backpropagation and gradient descent on the entire model can achieve the training goal.
[0029] When applying this model to a video to be processed, only the video frames on both sides of the frame interpolation moment and the input at that moment need to be taken and input into the model to obtain the frame at the frame interpolation moment.
[0030] Through meta-learning in optical flow estimation and frame interpolation, the present invention can achieve good frame interpolation with only one parameter forward pass. Experiments prove that compared with the existing technologies, the present invention can achieve frame interpolation at any moment and can obtain excellent visual effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is the overall process structure diagram used in the embodiment of the present invention.
[0032] Figure 2 It is the structure diagram of the residual dense block used in the embodiment of the present invention.
[0033] Figure 3 It is the structure diagram of the feature extractor used in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To make the above features and advantages of the present invention more obvious and understandable, specific embodiments are given below and will be described in detail in conjunction with the accompanying drawings. It should be noted that the specific number of layers, number of modules, number of functions, and settings for certain layers given in the following embodiments are only a preferred implementation manner and are not used for limitation. Those skilled in the art can select the number and set certain layers according to actual needs, which should be understandable.
[0035] This embodiment discloses a method for video frame interpolation at any time based on bidirectional meta-learning. Taking the video frame interpolation of the target content video at any time as an example, the specific description is as follows:
[0036] Step 1: Collect a large number of videos to form a content dataset.
[0037] Step 2: Build an optical flow estimation network model and a video frame interpolation network model.
[0038] The overall network structure is as Figure 1 shown.
[0039] The model is divided into two sub-networks: an optical flow estimator based on meta-learning and a video frame interpolation generator based on meta-learning.
[0040] Among them, the optical flow estimation part includes a preliminary optical flow estimator, a linear optical flow mixer, an optical flow kernel generator, a feature extractor, and a meta-optical flow refiner. The preliminary optical flow estimator is a UNet, where the resolution first decreases layer by layer through convolution and then is upsampled by transposed convolution. There are skip connections between the same resolutions. The linear optical flow mixer is a traditional algorithm, and the expression of the optical flow is:
[0041]
[0042]
[0043] The optical flow kernel generator is a residual dense network, which consists of 4 residual dense blocks. Each residual dense block includes 4 convolutional layers, and each convolutional layer has 32 channels. The feature extractor is also a residual dense network, which consists of 12 residual dense blocks. Each residual dense block includes 8 convolutional layers, and each convolutional layer has 64 channels. The structure of the residual dense block is as Figure 2 shown. The meta-optical flow refiner consists of 2 convolutional layers and 2 fully connected layers.
[0044] Among them, the video frame interpolation part includes a frame initial generator, a frame kernel generator, a feature extractor, and a meta-frame regressor. The frame initial generator is a traditional algorithm that directly uses bilinear interpolation to generate using optical flow and the corresponding adjacent frame images. The frame kernel generator is a residual dense network composed of 4 residual dense blocks. Each residual dense block includes 4 convolutional layers, and each convolutional layer has 32 channels. The feature extractor is also a residual dense network composed of 12 residual dense blocks. Each residual dense block includes 8 convolutional layers, and each convolutional layer has 64 channels. The meta-frame regressor is composed of 2 convolutional layers and 2 fully connected layers.
[0045] The connection method inside each device is sequential connection in turn, without parallelism, and the output result of the previous part will be directly input to the next part. The connection order is consistent with the above description order. For the fully connected layer directly connected after the convolutional layer, in order to enable the output of the convolutional layer to be input to the fully connected layer, the output of the convolutional layer will be flattened and dimension-reduced.
[0046] Step 3: Train the optical flow estimation and video frame interpolation network model.
[0047] The total loss function term of the model is:
[0048] L = λ r L r + λ f L f + λ w L w + λ s L s ,
[0049] In the formula, λ r , λ f , λ w , λ s are weight terms. Usually, λ r is set to 1, λ f is set to 0.02, λ w is set to 0.2, λ s is set to 0.5.
[0050] L r is the reconstruction loss function term:
[0051]
[0052] In the formula, is the reconstructed frame, and I t+α is the real frame at this moment.
[0053] L f is the refined optical flow regularization term:
[0054]
[0055] In the formula, the superscript r represents the refined result, and the superscript p represents the preliminary result.
[0056] L w is the reconstruction frame consistency regularization term:
[0057]
[0058] In the formula, W is the aforementioned frame preliminary generator.
[0059] L s is the smoothing loss term:
[0060]
[0061] In the formula, Δ is the gradient operator.
[0062] Step 4: In the inference stage, input two existing adjacent frames and specify the time of the frame to be interpolated, and finally output the interpolated frame result map at that time.
[0063] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the spirit and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.
Claims
1. A video frame interpolation method at any time based on bidirectional meta-learning, the steps of which include: Training a video frame interpolation network model using the selected training data set; Wherein, the video frames at two non-adjacent times in the training data set are used as the input of the video frame interpolation network model, and the video frame at any one of the two non-adjacent times is output, and training is carried out with the real video frame at this time as the supervision; For a video to be processed, determine the frame interpolation time of the video, and then take the video frames on both sides of the frame interpolation time in the video and the frame interpolation time and input them into the trained video frame interpolation network model to obtain the video frame at the frame interpolation time; Wherein, the video frame interpolation network model includes an optical flow estimator based on meta-learning and a video frame interpolation generator; among them, the optical flow estimator based on meta-learning includes a preliminary optical flow estimator, a linear optical flow mixer, an optical flow kernel generator, a feature extractor, a weight estimator and a meta-optical flow refiner; the video frame interpolation generator based on meta-learning includes an optical flow-based frame preliminary generator, a frame kernel generator, a feature extractor and a meta-frame regressor; the method for training the video frame interpolation network model is: Input the video frames at two non-adjacent times in the training data set into the optical flow estimator based on meta-learning and the video frame interpolation generator respectively; the processing steps of the optical flow estimator based on meta-learning for the input video frames at two non-adjacent times are as follows: a) The video frames I of two non-adjacent moments are t ,I t+1 Input into the preliminary optical flow estimator to preliminarily estimate the optical flow from the previous frame to the next frame and from the next frame to the previous frame respectively; b) For video frame I t and I t+1 are respectively downsampled to obtain which are input into the preliminary optical flow estimator to respectively estimate the optical flow from the previous frame to the next frame and from the next frame to the previous frame in the case of downsampling, and the estimation results are upsampled back to the original resolution, denoted as c) Based on the result obtained in step a), the first linear optical flow mixer calculates the optical flow corresponding to the moments of frame I from any moment α of the frame to be interpolated by using the method of linear optical flow mixing t and I t+1 , which are respectively denoted as The second linear optical flow mixer calculates, according to the optical flow corresponding to the moments of frame I from any moment of the frame to be interpolated by using the method of linear optical flow mixing t and I t+1 , which are respectively denoted as d) Input the specific time value t + α of the frame to be interpolated jointly into the first weight estimator, and then input the result output by the first weight estimator into the optical flow kernel generator to generate an optical flow kernel and input it into the meta-optical flow refiner; Input into the first feature extractor to obtain the abstract features of the optical flow and input it into the meta-optical flow refiner; e) The optical flow refiner performs a convolution operation on the optical flow kernel using the abstract features of the optical flow to obtain a refined optical flow Calculate the refined optical flow regularization term L based on the initially estimated optical flow and the refined optical flow f , Calculate the smoothing loss term L based on the refined optical flow s ; The processing steps of the video frame interpolation generator based on meta-learning for the input video frames at two non-adjacent times are as follows: Put I t and I t+1 into the input frame preliminary generator to obtain a frame at any time α of the frame to be interpolated initially generated According to calculate the reconstruction frame consistency regularization term L w ; The refined optical flow is jointly input into the second weight estimator together with the specific time value t+α of the frame to be interpolated, and the result output by the second weight estimator is input into the frame kernel generator to generate a frame kernel Frame Combine two known frames I t , I t+1 Input into the second feature extractor to obtain abstract features The meta-frame regressor is based on the abstract features performs convolution on the frame kernel to obtain a refined frame Based on the refined frame and the corresponding ground truth frame, calculate the reconstruction loss function term L r ; The reconstruction loss function term L r , the reconstruction frame consistency regularization term L w , the refined optical flow regularization term L f , the smoothing loss term L s are weighted and combined to optimize the video frame interpolation network model.
2. The method according to claim 1, wherein The total loss function term of the video frame interpolation network model is: L = λ r L r + λ f L f + λ w L w + λ s L s ; where λ r 、λ f 、λ w 、λ s are weight terms; I t+α For the corresponding true frame; W is a frame initial generator; Δ is a gradient operator.
3. The method according to claim 1 or 2, characterized in that, The feature extractor is a residual dense network, including 12 sequentially connected residual dense blocks; each residual dense block contains 8 sequentially connected convolutional layers, and each convolutional layer contains 64 channels.
4. The method according to claim 1 or 2, characterized in that, The kernel generator is a residual dense network, including 4 sequentially connected residual dense blocks; each residual dense block contains 4 sequentially connected convolutional layers, and each convolutional layer contains 32 channels.
5. A video frame interpolation system at any time based on bidirectional meta-learning, characterized in that, Include an optical flow estimator based on meta-learning and a video frame interpolation generator; among them, the optical flow estimator based on meta-learning includes a preliminary optical flow estimator, a linear optical flow mixer, an optical flow kernel generator, a feature extractor, a weight estimator and a meta-optical flow refiner; the video frame interpolation generator based on meta-learning includes an optical flow-based frame preliminary generator, a frame kernel generator, a feature extractor and a meta-frame regressor; The preliminary optical flow estimator is used to estimate the video frames I at two non - adjacent moments of the input t and I t+1 to obtain the optical flow from the previous frame to the next frame and the optical flow from the next frame to the previous frame, and input them into the linear optical flow mixer; and estimate the t video frames I t+1 respectively downsampled to obtain respectively, to obtain the optical flow from the previous frame to the next frame and the optical flow from the next frame to the previous frame, and upsample the estimation results back to the original resolution, denoted as The linear optical flow mixer is used to calculate the optical flow at the corresponding times from any time α of the to-be-inserted frame to frame I t and I t+1 respectively denoted as and according to calculate the optical flow at the corresponding times from any time of the to-be-inserted frame to frame I t and I t+1 respectively denoted as The weight estimator is used to estimate weights according to the input and the specific time value t+α of the frame to be interpolated, and input the estimated weights into the optical flow kernel generator to generate an optical flow kernel The first feature extractor is used to, according to the input obtain the abstract features of the optical flow The above-mentioned meta-optical flow refiner is used to perform a convolution operation on the optical flow kernel by using the abstract features of the optical flow to obtain a refined optical flow The frame preliminary generator is used to generate a frame at any moment according to the input I t 、I t+1 and The refined optical flow and the specific time value t+α of the frame to be interpolated are jointly input into the weight estimator, and then the result output by the weight estimator is input into the frame kernel generator to generate a frame kernel A second feature extractor, configured to, based on an input frame and frame I t 、I t+1 obtain an abstract feature The meta-frame regressor is configured to, based on the abstract feature perform convolution on a frame kernel to obtain a refined frame 6. A server, characterized in that, Include a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Video frame interpolation method, model training method, and corresponding device
WO2022033048A1