Method for training high-resolution video enhancement model based on continuous motion
By introducing a motion prior extractor and correction network into the video interpolation model, combining the random interpolation time training strategy and multiple loss functions, the problem of insufficient interpolation effect in the existing technology in high-resolution video and large-scale motion scenarios is solved, and a more efficient and accurate video interpolation effect is achieved.
Patent Information
- Application Number
- CN202510057811.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-14
Smart Images

Figure CN120032290A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of artificial intelligence and image processing, and in particular relates to a training method for a high-resolution video enhancement model based on continuous motion. Background Art
[0002] Video Frame Interpolation (VFI) is an important video processing technology, which aims to generate missing intermediate frames between existing video frames, thereby improving the temporal resolution of the video and enhancing the viewing experience. Video interpolation can not only be used for ordinary video enhancement, but is also widely used in slow motion generation, video compression, image restoration and animation production. In recent years, with the rapid development of deep learning technology, video interpolation methods have been continuously innovated, especially in high-resolution videos and large-scale motion scenes. Some progress has been made. However, although the existing technology has good performance in theory, it still faces many challenges in practical applications, especially in high-resolution videos, large-scale motion and complex scenes. The interpolation effect is far from expected.
[0003] In recent years, interpolation methods based on optical flow have been widely used, especially when dealing with small motion or stable video content. For example, Xin Jin et al. proposed a pyramid-structured recurrent network method in the paper "A unified pyramid recurrent network for video frame interpolation", which generates intermediate frames through optical flow estimation. This method guides the interpolation process by calculating the optical flow using the correlation between adjacent frames. However, the main drawback of this method is that the calculation of optical flow depends on correlation matching and lacks sufficient interpretability. Especially in the face of high-resolution videos and large-scale motion, it is often unable to effectively capture continuous motion trajectories. Due to the instability of motion and the error of optical flow estimation, the generated intermediate frames appear blurred and distorted.
[0004] In order to solve the interpolation problem in large motion scenes, Chunxu Liu et al. proposed a sparse global matching method in "Sparse global matching for video frame interpolation with large motion". This method attempts to enhance the processing capability of large motion and high-resolution video interpolation by introducing additional data sets. This method extracts global motion information through sparse matching and synthesizes intermediate frames through global information. However, although this method provides an effective strategy for solving large motion scenes in theory, there are still several problems in practical applications. First, this method introduces new data sets during training, but its training methods and strategies still lack sufficient systematic exploration, resulting in its performance in high-resolution video interpolation not as expected. Secondly, the global attention mechanism it adopts causes the computational complexity to increase quadratically with the video resolution, which makes its computational requirements increase sharply when processing high-resolution videos. In addition, in scenes with drastic dynamic changes, the local limitations of the sparse matching method still limit its performance in processing complex scenes and cannot fully meet the needs of high-resolution and large-scale motion interpolation.
[0005] In addition, Guozhen Zhang et al. proposed a state space model-based interpolation method in "VFIMamba: Video Frame Interpolation with State Space Models", which mainly explored the training method of fixed-time interpolation. By using the state space model, the model can effectively transfer information between frames to generate accurate intermediate frames. Although this method performs well in fixed interpolation scenarios, its insufficient exploration of multi-frame insertion (i.e., generating multiple intermediate frames) limits its application in complex dynamic scenes. For example, in high-speed motion and rapidly changing scenes, the fixed interpolation method has difficulty capturing the details of continuous motion, resulting in distortion and blurring of the generated intermediate frames.
[0006] An important challenge facing high-resolution video interpolation is how to maintain the accuracy and efficiency of the interpolation process without introducing too much computational overhead. Many existing methods often require a lot of computing resources and complex model structures when interpolating high-resolution videos. When dealing with large motion scenes, how to accurately model the motion trajectory and ensure the details and clarity of the interpolated image is a major bottleneck of current technology. The processing of high-resolution videos requires higher detail retention capabilities, but most existing methods ignore how to efficiently handle complex motion under high-resolution conditions, resulting in frequent distortion in detail and texture retention.
[0007] Although existing video interpolation methods have achieved good results under certain conditions, they generally have the following problems: First, how to accurately model motion trajectories in complex scenes, especially in the case of high resolution and large-scale motion, is still an unsolved problem. Second, most existing methods focus on a specific type of scene or task, and lack broad adaptability to a variety of scenes and tasks. Especially when faced with videos with drastic dynamic changes, existing interpolation methods often fail to accurately capture motion details, resulting in blurred and distorted generated images. Finally, how to effectively conduct efficient training and reduce computational overhead remains a research focus. Summary of the invention
[0008] In order to solve the above problems existing in the prior art, the present invention provides a training method for a high-resolution video enhancement model based on continuous motion and a video frame insertion method. The technical problem to be solved by the present invention is achieved through the following technical solutions:
[0009] In a first aspect, an embodiment of the present invention provides a method for training a high-resolution video enhancement model based on continuous motion, the method comprising:
[0010] Constructing a video interpolation model; wherein the video interpolation model includes a feature extractor, a motion prior extractor, an optical flow estimator and a correction network; the feature extractor is used to perform multi-dimensional feature extraction on two input frames; the motion prior extractor is used to obtain displacement information of the intermediate frame based on the high-dimensional features, correlation measurement and interpolation time; the optical flow estimator is used to obtain optical flow estimation information based on the displacement information of the intermediate frame, thereby obtaining a preliminary intermediate frame image; the correction network is used to correct the preliminary intermediate frame image according to the multi-dimensional features, and output a predicted intermediate frame image;
[0011] Based on the first video data set, the video frame insertion model is pre-trained by adopting a training strategy with random frame insertion time and a preset multiple loss function; wherein the multiple loss function is constructed based on a Laplace loss function and an optical flow loss function;
[0012] Based on a second video data set containing long-term high-resolution video clips, the pre-trained video interpolation model is optimized and trained to obtain a trained video interpolation model; the trained video interpolation model is used to output an interpolated intermediate frame image for two input frame images.
[0013] In one embodiment of the present invention, the feature extractor includes three CNN modules and two Transformer modules; wherein the two Transformer modules are used to output high-dimensional features.
[0014] In one embodiment of the present invention, the process of the motion prior extractor obtaining the displacement information of the intermediate frame based on the correlation metric and the interpolation time according to the high-dimensional features output by the feature extractor includes:
[0015] Using two motion prior extraction modules of the motion prior extractor, high-dimensional features of two frames of images are received; wherein the two frames of images are respectively a first frame of image and a second frame of image;
[0016] For each pixel position, according to the high-dimensional features of the two frames of images, calculate a correlation value between the pixel value of the pixel position in the first frame of image and each neighboring pixel value of the pixel position in the second frame of image;
[0017] The correlation value obtained at each pixel position is processed by Softmax;
[0018] According to the correlation value after Softmax processing and the coordinates of each pixel position, the displacement information of each pixel is obtained;
[0019] The displacement information of the corresponding pixel in the intermediate frame is obtained according to the product of the displacement information of each pixel and the interpolation time.
[0020] In one embodiment of the present invention, the displacement information of each pixel is obtained according to the correlation value after Softmax processing and the coordinates of each pixel position, including:
[0021] Perform tensor transformation on the correlation value after Softmax processing;
[0022] For each pixel position, after mapping its coordinates from two dimensions to the same dimension as the feature, the current correlation value of the pixel position is multiplied by its coordinates after dimension conversion, and the results are summed within the neighborhood to obtain the displacement information of each pixel.
[0023] In one embodiment of the present invention, the optical flow estimator obtains the optical flow estimation information according to the displacement information of the intermediate frame, and thus obtains the preliminary intermediate frame image using the formula:
[0024] Iwarp=M.×warp(I 1 ,Ft →1 )+(1-M).×warp(I 2 ,Ft →2 );
[0025] Among them, Iwarp is the preliminary intermediate frame image; M is the mask generated when generating the optical flow, which is used to deal with the occlusion in the interpolation process; .× represents the dot product; I 1 and I 2are the first frame image and the second frame image of the two frames respectively; t is the interpolation time; Ft →1 is the optical flow estimation from the intermediate frame to the first frame; Ft →2 is the optical flow estimation from the middle frame to the second frame image; Ft →1 , Ft →2 and M constitute the optical flow estimation information; warp(·) is the operation of distorting the image using the optical flow.
[0026] In one embodiment of the present invention, the correction network corrects the preliminary intermediate frame image according to the multi-dimensional features and outputs a predicted intermediate frame image, comprising:
[0027] The correction network obtains the residual of the preliminary intermediate frame image according to the multi-dimensional features;
[0028] The preliminary intermediate frame image and the residual are summed to obtain a predicted intermediate frame image.
[0029] In one embodiment of the present invention, the first video dataset is a Vimeo90K dataset, and the second video dataset is a LAVIB dataset.
[0030] In one embodiment of the present invention, before pre-training the video frame insertion model using a training strategy with random frame insertion time and a preset multiple loss function, the method further includes:
[0031] Performing data enhancement processing on the first video data set and setting training parameters;
[0032] Before optimizing and training the pre-trained video frame insertion model, the method further includes:
[0033] The data enhancement processing is performed on the second video data set, and training parameters are set.
[0034] In one embodiment of the present invention, the multiple loss function is a weighted sum of a Laplace loss function and an optical flow loss function.
[0035] In a second aspect, an embodiment of the present invention provides a video frame insertion method, the video frame insertion method comprising:
[0036] Obtain the first frame image and the second frame image to be inserted;
[0037] The first frame image, the second frame image and the interpolation time are input into a pre-trained video interpolation model, and an inserted intermediate frame image is output; wherein the pre-trained video interpolation model is obtained according to the training method of the high-resolution video enhancement model based on continuous motion described in the first aspect.
[0038] The training method of the high-resolution video enhancement model based on continuous motion provided by the present invention has the following beneficial effects:
[0039] First, compared with the existing video interpolation model, the model proposed in the present invention performs continuous motion modeling through Correlation, ensuring the accuracy of motion information of intermediate frames, especially in large motion and complex scenes.
[0040] Second, through the pre-training strategy of the present invention, the model has the ability to model continuous motion, which can better solve the video frame insertion scene with large motion amplitude. After further fine-tuning, the model's effect in processing high-resolution and large-span videos is further improved.
[0041] Third, the model can not only handle video interpolation tasks with long time spans, but also maintain smooth and natural intermediate frame effects in high-resolution videos with large motion amplitudes, which is suitable for various practical application scenarios.
[0042] The present invention can effectively eliminate the uncertainty of motion between different frames by modeling continuous motion between frames, thereby realizing accurate motion estimation of specific positions between frames by the model. Furthermore, based on accurate motion estimation, the present invention can achieve precise optical flow estimation and intermediate frame generation. This method is used to improve the accuracy and effect of high-resolution video interpolation, especially when dealing with large-scale motion and complex scenes, with stronger robustness and accuracy. The trained video interpolation model can be used in entertainment, digital media, and social public security, such as animation production, video acquisition, etc., to achieve accurate interpolation effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A schematic flow chart of a training method for a high-resolution video enhancement model based on continuous motion provided by an embodiment of the present invention;
[0044] Figure 2 A schematic diagram of the structure and principle of a video frame insertion model provided by an embodiment of the present invention;
[0045] Figure 3 A schematic diagram of a flow chart of a video frame insertion method provided by an embodiment of the present invention;
[0046] Figure 4 This is the result of an embodiment of the present invention on a large-span video interpolation dataset;
[0047] Figure 5 The results of the embodiments of the present invention on high-resolution videos and large-motion videos;
[0048] Figure 6 These are the results of more large-span video interpolation test sets according to the embodiments of the present invention. DETAILED DESCRIPTION
[0049] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0050] In order to solve the problem of inter-frame motion uncertainty faced by existing high-resolution video interpolation methods, an embodiment of the present invention provides a training method for a high-resolution video enhancement model based on continuous motion and a video interpolation method.
[0051] In a first aspect, an embodiment of the present invention provides a training method for a high-resolution video enhancement model based on continuous motion, such as Figure 1 As shown, the method may include the following steps S1 to S3:
[0052] S1, build a video frame insertion model;
[0053] The video interpolation model includes a feature extractor, a motion prior extractor, an optical flow estimator and a correction network; the feature extractor is used to perform multi-dimensional feature extraction on two input frames; the motion prior extractor is used to obtain displacement information of the intermediate frame based on the high-dimensional features, correlation measurement and interpolation time; the optical flow estimator is used to obtain optical flow estimation information based on the displacement information of the intermediate frame, so as to obtain a preliminary intermediate frame image; the correction network is used to correct the preliminary intermediate frame image according to the multi-dimensional features, and output a predicted intermediate frame image;
[0054] For the structure of the video interpolation model, see Figure 2 The video interpolation model of the embodiment of the present invention is improved and designed based on the overall framework of the model in the document "VFIMamba: Video Frame Interpolation with State Space Models". The model in the document includes a feature extractor, an optical flow estimator and a correction network. The present invention adds a motion prior extractor on the basis of the model to extract motion information to ensure the accuracy of the estimated optical flow.
[0055] The following describes each part of the video frame insertion model separately.
[0056] ①Feature Extractor
[0057] The feature extractor is used to extract multi-dimensional features from the two input frames of images. For ease of understanding, the two input frames of images are represented by I 1 and I 2 Refer to Figure 2 , the feature extractor includes three CNN modules and two Transformer modules, then the feature extractor can1 and I 2 , extracting features of five dimensions, wherein the two Transformer modules are used to output high-dimensional features.
[0058] ②Motion prior extractor
[0059] The motion prior extractor is used to obtain the displacement information of the intermediate frame according to the high-dimensional features based on the correlation metric and the interpolation time. The process may include the following steps:
[0060] Step A1, using two motion prior extraction modules of the motion prior extractor to receive high-dimensional features of two frames of images;
[0061] from Figure 2 It can be seen that the motion prior extractor includes two motion prior extraction modules;
[0062] The two frames of images are respectively the first frame of image and the second frame of image. 1 and I 2 It is understandable that I 1 and I 2 The sizes are consistent and the pixel positions are one-to-one corresponding.
[0063] Step A2, for each pixel position, calculating a correlation value between a pixel value of the pixel position in the first frame image and each neighboring pixel value of the pixel position in the second frame image according to high-dimensional features of the two frames of images;
[0064] In video processing, motion information between frames is crucial for generating continuous and natural video interpolation. The present invention ensures that the model can effectively capture the motion pattern in the video by extracting continuous motion priors, thereby generating smooth intermediate frames. First, the displacement information of each pixel in the image is estimated by calculating the correlation between frames. Correlation is a common method to measure the similarity between two frames of images, which can help identify which pixels have moved between the two frames, and then estimate the displacement. The calculation method is shown in Formula 1:
[0065] Correlation(I 1 ,I 2 ) xymn =I 1 (x,y)I 2 (m,n) (1);
[0066] Among them, I(x,y) represents the pixel value of the corresponding position, such as I 1 (x,y) represents I 1The pixel value at position (x, y) in the equation is (m, n), which represents the position of the pixel in the neighborhood of (x, y). Usually, the neighborhood is an area of size (k×k), where k can be 3, etc. 1 (x,y), we can use I 1 and I 2 The high-dimensional features, and I 2 A correlation value is calculated for each pixel in the (m,n) neighborhood of (x,y).
[0067] Let I 1 and I 2 The length and width of the image are h and w respectively, that is, the image contains h×w pixels, that is, there are h×w positions (x, y), I 1 Each pixel position in can be calculated by the above correlation and I 2 Each k×k neighborhood position in the image gets a correlation value, so Correlation(I 1 ,I 2 ) is of size (h,w,k,k).
[0068] Step A3, performing Softmax processing on the correlation value obtained at each pixel position;
[0069] In order to make the displacement estimation more accurate, the present invention performs Softmax processing on the correlation value obtained at each pixel position to enhance the matching degree at the pixel level. When calculating, it is first transformed into a tensor of size (h,w,k×k), and the calculation is shown in Formula 2:
[0070]
[0071] Where i represents the serial number of the pixel position;
[0072] Softmax can give more relevant pixel displacements greater weights.
[0073] Step A4, obtaining the displacement information of each pixel according to the correlation value after Softmax processing and the coordinates of each pixel position;
[0074] Specifically, step A4 may include:
[0075] Step A41, performing tensor transformation processing on the correlation value after Softmax processing;
[0076] This step is to convert Correlation(I 1 ,I 2 ) back to a tensor of size (h,w,k,k).
[0077] Step A42, for each pixel position, after mapping its coordinates from two dimensions to the same dimension as the feature, multiply the current correlation value of the pixel position by its coordinates after dimension conversion, and sum the results within the neighborhood to obtain the displacement information of each pixel.
[0078] Since the motion prior extractor extracts motion information at the feature level, the coordinates used to calculate the displacement information will first be mapped from two dimensions to the same dimension as the feature through a linear mapping layer, and then the displacement information will be calculated using Formula 3.
[0079]
[0080] In Formula 3, (x, y) is the pixel position coordinate after mapping from two dimensions to the same dimension as the feature. Δx(x, y) represents the displacement information of the corresponding pixel.
[0081] Step A5, obtaining the displacement information of the corresponding pixel of the intermediate frame according to the product of the displacement information of each pixel and the interpolation time.
[0082] Obtaining the displacement information of each pixel means obtaining the displacement information between frames, and then adjusting the displacement information according to the interpolation time.
[0083] Assuming that the interpolation time is t (a value between 0 and 1), the displacement information of the corresponding pixel finally inserted is the product of the displacement information of the pixel obtained in step A4 and the interpolation time, as shown in Formula 4:
[0084] Δxmid = Δx·t (4);
[0085] Wherein, Δx is the displacement information of the pixel (x, y) obtained in step A4, and Δxmid represents the displacement information of the corresponding pixel in the intermediate frame.
[0086] Through this process, the embodiment of the present invention can generate and model different priors for different interpolation frame times t.
[0087] ③Optical flow estimator
[0088] Based on the above prior, we can use this prior to calculate the intermediate frame optical flow Ft →1 , Ft →2 Compared with the previous method of estimating the optical flow of the intermediate frame using features, the use of this prior for the estimation of the optical flow of the intermediate frame has stronger interpretability and stronger robustness when performing multi-frame insertion. The intermediate frame can be preliminarily obtained through the optical flow of the intermediate frame.
[0089] The optical flow estimator obtains the optical flow estimation information according to the displacement information of the intermediate frame, so as to obtain the preliminary intermediate frame image using the formula:
[0090] Iwarp = M × warp(I 1 , Ft →1 ) + (1 - M) × warp(I 2 , Ft →2 ) (5);
[0091] Among them, Iwarp is the preliminary intermediate frame image; M is the mask generated simultaneously when generating the optical flow, which is used to process the occluder during the frame interpolation process; × represents dot product; I 1 and I 2 are respectively the first frame image and the second frame image in the two frame images; t is the frame interpolation time; Ft →1 is the optical flow estimation from the intermediate frame to the first frame image; Ft →2 is the optical flow estimation from the intermediate frame to the second frame image; Ft →1 , Ft →2 and M constitute the optical flow estimation information, which is represented by F in Figure 2 ; warp(·) is the operation of warping the image using the optical flow.
[0092] ④ Correction network
[0093] The process that the correction network corrects the preliminary intermediate frame image according to the multi-dimensional features and outputs the predicted intermediate frame image includes:
[0094] 1) The correction network obtains the residual of the preliminary intermediate frame image according to the multi-dimensional features;
[0095] Please refer to Figure 2 . The correction network includes multiple CNN modules, and the first 5 CNN modules each receive one of the multi-dimensional features. Through the processing of multiple CNN modules of the correction network, the residual ΔI of the preliminary intermediate frame image is output.
[0096] 2) Perform summation processing on the preliminary intermediate frame image and the residual to obtain the predicted intermediate frame image.
[0097] Then, perform summation processing on the preliminary intermediate frame image Iwarp and the residual ΔI. The summation processing is shown by the adder in Figure 2 to obtain the predicted intermediate frame image Ipred.
[0098] For the specific structures of the feature extractor, optical flow estimator, and correction network, please refer to the relevant literature for understanding.
[0099] S2. Based on the first video dataset, use the training strategy with random frame interpolation time and the preset multiple loss function to pre-train the video frame interpolation model; among them, the multiple loss function is constructed based on the Laplace loss function and the optical flow loss function;
[0100] In S1, a video interpolation model is constructed. In S2, the present invention proposes a general pre-training strategy, which enables the model to have stronger continuous motion modeling capabilities during the training phase. In order to effectively handle the uncertainty of inter-frame displacement in large motion and high-resolution videos, the first video dataset is selected as the Vimeo90K dataset, and multiple loss functions are designed to optimize the motion modeling of the model.
[0101] Specifically, we first introduce Laplace Loss to optimize the pixel difference between the generated intermediate frame and the real frame. The formula is:
[0102]
[0103] Where Ipred is the intermediate frame generated by the model, and Igt is the real intermediate frame. L represents the average pooling operation, i represents the number of average pooling times, and N represents the maximum number of average pooling times. The value of N can be set as needed, for example, 5.
[0104] The role of the Laplace loss function is to maintain the structural consistency between the real frame and the predicted frame.
[0105] In addition, the present invention also designs a loss function based on optical flow to further improve the quality of generated frames by optimizing the motion consistency between frames. For the displacement of each pixel, the model needs to infer the direction of movement through optical flow and minimize the difference between the predicted displacement and the actual displacement. The optical flow loss function is as follows:
[0106]
[0107] Optical flow loss further enhances the model's motion modeling capabilities, especially in large motion or complex scenes, and is able to maintain motion consistency, resulting in a more natural interpolation effect.
[0108] The multiple loss function is a weighted sum of the Laplace loss function and the optical flow loss function, and the weights are set as needed.
[0109] Optionally, before pre-training the video interpolation model by adopting a training strategy with random interpolation time and a preset multiple loss function, the training method of the high-resolution video enhancement model based on continuous motion further includes:
[0110] Performing data enhancement processing on the first video data set and setting training parameters;
[0111] Specifically, during the training process, the Vimeo90K dataset is used for data augmentation. During each training, three frames of video are randomly selected, and 256×256 image blocks are cropped from the three frames of video. Data augmentation such as random flipping, time reversal, and random rotation is performed. The images of the two frames before and after are used as the two frames to be interpolated, and the middle frame is used as the real middle frame, i.e., the label data. The training parameters are set as follows: the training batch size is set to 32, the optimizer uses AdamW, the training cycle is set to 300, and the learning rate is increased from 1e- through the cosine annealing strategy. 4 Gradually decays to 1e- 5 .
[0112] Compared with the previous method of directly using three consecutive frames for interpolation training at time t=0.5, the pre-training strategy of the present invention uses random interpolation time, which can enable the model to have the ability to model continuous motion and enhance the performance of the model for high-resolution videos and large-motion videos. The improvement effect of the pre-training method proposed by the present invention can be reflected in Table 4 of the experiment below.
[0113] The specific training process can be understood by referring to the conventional neural network training process, which will not be described in detail here.
[0114] S3, based on a second video data set containing long high-resolution video clips, optimizing and training the pre-trained video interpolation model to obtain a trained video interpolation model; the trained video interpolation model is used to output an interpolated intermediate frame image for two input frames of images.
[0115] In order to further improve the model's interpolation capability over a long time span, the present invention uses long videos and high-resolution datasets for model fine-tuning. The second video dataset is the LAVIB (Large-scale Video Interpolation Benchmark) dataset, and its reference is "Alexandros Stergiou. Lavib: Large-scale video interpolation benchmark. In NeurIPS, 2024". The LAVIB dataset contains video clips with a longer time span. The present invention extracts a continuous 20-frame segment from each video clip and randomly selects three frames from them for training.
[0116] Similarly, before optimizing the pre-trained video frame insertion model, the method further includes:
[0117] The data enhancement processing is performed on the second video data set, and training parameters are set.
[0118] Specifically, each frame is processed by data augmentation operations such as random scaling and cropping. Each frame is scaled to 256×256, the training batch size is set to 16, and the learning rate is set from 1e- 5 starts and decays to 1e- 6 The training settings in the fine-tuning phase are the same as those in pre-training. By fine-tuning on longer video clips, the model can capture motion information over a longer time span.
[0119] The definition of the loss function is consistent with the pre-training stage, and the goal is still to optimize the model's inter-frame interpolation ability by minimizing the weighted sum of the Laplace loss and the optical flow loss. Fine-tuning enables the model to handle more complex video interpolation tasks, especially in high-resolution video interpolation, and to generate high-quality intermediate frames that are smoother and more natural.
[0120] Compared with the previous method, the performance of the model in large motion range continuous modeling can be further improved by fine-tuning the model specifically on long videos with high resolution, which can be understood from the results in Table 1 in the following experiment. And by fine-tuning, the amount of computation required for model training is greatly reduced compared to directly training on videos with large motion range. In this way, a better effect is achieved on the large motion range dataset.
[0121] The training method of the high-resolution video enhancement model based on continuous motion provided by the embodiment of the present invention has the following beneficial effects:
[0122] First, compared with the existing video interpolation model, the model proposed in the present invention performs continuous motion modeling through Correlation, ensuring the accuracy of motion information of intermediate frames, especially in large motion and complex scenes.
[0123] Second, through the pre-training strategy of the present invention, the model has the ability to model continuous motion, which can better solve the video frame insertion scene with large motion amplitude. After further fine-tuning, the model's effect in processing high-resolution and large-span videos is further improved.
[0124] Third, the model can not only handle video interpolation tasks with long time spans, but also maintain smooth and natural intermediate frame effects in high-resolution videos with large motion amplitudes, which is suitable for various practical application scenarios.
[0125] The present invention can effectively eliminate the uncertainty of motion between different frames by modeling continuous motion between frames, thereby realizing accurate motion estimation of specific positions between frames by the model. Furthermore, based on accurate motion estimation, the present invention can realize precise optical flow estimation and intermediate frame generation. This method can improve the accuracy and effect of high-resolution video interpolation, especially when dealing with large-scale motion and complex scenes, with stronger robustness and accuracy. The trained video interpolation model can be used in entertainment, digital media, social public security and other fields, such as animation production, video acquisition, etc.
[0126] In a second aspect, corresponding to the embodiment of the training method of the high-resolution video enhancement model based on continuous motion described in the first aspect, the embodiment of the present invention further provides a video frame insertion method, such as Figure 3 As shown, the video frame insertion method includes:
[0127] S100, obtaining a first frame image and a second frame image to be inserted;
[0128] S200, input the first frame image, the second frame image and the interpolation time into a pre-trained video interpolation model, and output an inserted intermediate frame image; wherein the pre-trained video interpolation model is obtained according to the training method of the high-resolution video enhancement model based on continuous motion described in the first aspect.
[0129] The present invention can output high-resolution intermediate frames for low-resolution videos input by a pre-trained and fine-tuned video interpolation model. These intermediate frames are not only visually similar to real frames, but also highly consistent in motion information.
[0130] Regarding the structure and processing process of the pre-trained video interpolation model, please refer to the relevant content of the first aspect and will not be elaborated here.
[0131] The video interpolation method provided in the embodiment of the present invention is completed by using a pre-trained video interpolation model, can accurately estimate the motion of specific positions between frames, can achieve precise optical flow estimation and intermediate frame generation, and can improve the accuracy and effect of high-resolution video interpolation, especially when processing large-scale motion and complex scenes, with stronger robustness and accuracy.
[0132] 1. Simulation conditions
[0133] The experiment of the present invention was carried out on an Intel(R) Xeon(R) CPU E5-2620 v4@2.10GHz, an NVIDIARTX 2080Ti, and an Ubuntu 16.04 operating system. The programming language used was Python, the deep learning network framework used was Pytorch, and the training data sets used were Vimeo90K and LAVIB.
[0134] In order to evaluate the performance of the method fairly and objectively, the experiment of the present invention ensures that all the comparison models are trained by the same training dataset. The experiment of the present invention uses the same settings to train other state-of-the-art algorithms for comparison. Specifically including UPRNET, VFIMamba, SGM-VFI, and AMT, their references are "Xin Jin, Longhai Wu, JieChen, Youxin Chen, Jayoon Koo, and Cheul-Hee Hahm. Aunified pyramid recurrent network for video frame interpolation. In 2023 IEEE / CVF Conference on ComputerVision and Pattern Recognition (CVPR), pages 1578-1587,2023","Guozhen Zhang,Chunxu Liu,Yutao Cui,Xiaotong Zhao,Kai Ma,and Limin Wang.VFIMamba: Video FrameInterpolation with State Space Models,2024.","Chunxu Liu,Guozhen Zhang,RuiZhao,and Limin Wang.Sparse global matching for video frame interpolation with large motion.In 2024IEEE / CVF Conference on Computer Vision and PatternRecognition(CVPR),pages 19125-19134,Los Alamitos, CA, USA, 2024" and "Zhen Li, Zuo-Liang Zhu, Ling-Hao Han, Qibin Hou, ChunLe Guo, and Ming-Ming Cheng. Amt: All-pairsmulti-field transforms for efficient frame interpolation. In 2023 IEEE / CVFConference on Computer Vision and Pattern Recognition (CVPR), pages9801-9810, 2023."In addition to the retrained algorithms, the experiments in this invention also compared the models provided by themselves and the models provided by some other algorithms including DQBC (IJCAI 2023) and RIFE (ECCV 2022).
[0135] 2. Simulation content
[0136] Table 1 Performance of each model on large-span high-resolution videos
[0137]
[0138]
[0139] In Table 1, “-RT” indicates the model trained according to the training setting proposed by the present invention, “V” indicates Vimeo90K, and “L” indicates LAVIB;
[0140] Table 2 Performance of each model on large motion videos
[0141]
[0142] In Table 2, "-RT" indicates the retrained model of the present invention, "-PT" indicates the pretrained model, "V" indicates Vimeo90K, and "L" indicates LAVIB.
[0143] Table 3 Performance of each model on high-resolution videos
[0144]
[0145]
[0146] In Table 3, "-RT" indicates the model retrained by the present invention, "-PT" indicates the model trained under the pre-training strategy, "V" indicates Vimeo90K, and "L" indicates LAVIB.
[0147] It can be seen from Tables 1, 2, and 3 that the method of the present invention is much better than the existing methods in objective evaluation indicators PSNR (peak signal-to-noise ratio) and SSIM (structural similarity), and the retrained method of the present invention shows greater superiority in large motion video interpolation and high-resolution video interpolation tasks compared with the weights provided by its authors. From the visual effect point of view, other existing methods may have serious distortion in a certain local position while the present invention can maintain its general shape. At the same time, in order to verify the superiority of the pre-training method proposed by the present invention, a comparison between the results of the pre-training algorithm and the best existing method is also given in Table 4. It can be seen that after pre-training of the present invention, the existing methods can surpass the best existing methods to a certain extent. In addition, please refer to the visual evaluation results of each model on large-span high-resolution videos. Figures 4 to 6 shown. Figure 4 To show the results on a large span video interpolation dataset, Figure 5 For high-resolution and large-motion video results, Figure 6 Results on more large span video interpolation test sets.
[0148] Table 4 Comparison of each model after pre-training with the best existing method VFIMamba
[0149]
[0150] In Table 4, "-FT" represents the model after fine-tuning of the present invention, "-PT" represents the model trained according to the pre-training strategy of the present invention, "V" represents Vimeo90K, and "L" represents LAVIB.
[0151] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification.
[0152] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A training method for a high-resolution video enhancement model based on continuous motion, characterized in that: include: Constructing a video interpolation model; wherein the video interpolation model includes a feature extractor, a motion prior extractor, an optical flow estimator and a correction network; the feature extractor is used to perform multi-dimensional feature extraction on two input frames; the motion prior extractor is used to obtain displacement information of the intermediate frame based on the extracted high-dimensional features, correlation measurement and interpolation time; the optical flow estimator is used to obtain optical flow estimation information based on the displacement information of the intermediate frame, thereby obtaining a preliminary intermediate frame image; the correction network is used to correct the preliminary intermediate frame image according to the multi-dimensional features, and output a predicted intermediate frame image; Based on the first video data set, the video frame insertion model is pre-trained by adopting a training strategy with random frame insertion time and a preset multiple loss function; wherein the multiple loss function is constructed based on a Laplace loss function and an optical flow loss function; Based on a second video data set containing long-term high-resolution video clips, the pre-trained video interpolation model is optimized and trained to obtain a trained video interpolation model; the trained video interpolation model is used to output an interpolated intermediate frame image for two input frame images.
2. The training method of the high-resolution video enhancement model based on continuous motion according to claim 1, characterized in that: The feature extractor includes three CNN modules and two Transformer modules; wherein the two Transformer modules are used to output high-dimensional features.
3. The training method of the high-resolution video enhancement model based on continuous motion according to claim 1, characterized in that: The process of obtaining the displacement information of the intermediate frame by the motion prior extractor based on the high-dimensional features output by the feature extractor and the correlation measurement and the interpolation time includes: Using two motion prior extraction modules of the motion prior extractor, high-dimensional features of two frames of images are received; wherein the two frames of images are respectively a first frame of image and a second frame of image; For each pixel position, according to the high-dimensional features of the two frames of images, calculate a correlation value between the pixel value of the pixel position in the first frame of image and each neighboring pixel value of the pixel position in the second frame of image; The correlation value obtained at each pixel position is processed by Softmax; According to the correlation value after Softmax processing and the coordinates of each pixel position, the displacement information of each pixel is obtained; The displacement information of the corresponding pixel in the intermediate frame is obtained according to the product of the displacement information of each pixel and the interpolation time.
4. The training method of the high-resolution video enhancement model based on continuous motion according to claim 3 is characterized in that: The displacement information of each pixel is obtained according to the correlation value after Softmax processing and the coordinates of each pixel position, including: Perform tensor transformation on the correlation value after Softmax processing; For each pixel position, after mapping its coordinates from two dimensions to the same dimension as the feature, the current correlation value of the pixel position is multiplied by its coordinates after dimension conversion, and the results are summed within the neighborhood to obtain the displacement information of each pixel.
5. The training method of the high-resolution video enhancement model based on continuous motion according to claim 1, characterized in that: The optical flow estimator obtains the optical flow estimation information according to the displacement information of the intermediate frame, so as to obtain the preliminary intermediate frame image using the formula: Iwarp=M.×warp(I1,Ft →1 )+(1-M).×warp(I2,Ft →2 ); Where Iwarp is the preliminary intermediate frame image; M is the mask generated when generating the optical flow, which is used to handle the occluders in the interpolation process; .× represents the dot product; I1 and I2 are the first and second frames of the two images respectively; t is the interpolation time; Ft →1 is the optical flow estimation from the intermediate frame to the first frame; Ft →2 is the optical flow estimation from the middle frame to the second frame image; Ft →1 , Ft →2 and M constitute the optical flow estimation information; warp(·) is the operation of distorting the image using the optical flow.
6. The training method of the continuous motion based high-resolution video enhancement model according to claim 1, characterized in that: The process of the correction network correcting the preliminary intermediate frame image according to the multi-dimensional features and outputting a predicted intermediate frame image includes: The correction network obtains the residual of the preliminary intermediate frame image according to the multi-dimensional features; The preliminary intermediate frame image and the residual are summed to obtain a predicted intermediate frame image.
7. The training method of the high-resolution video enhancement model based on continuous motion according to claim 1, characterized in that: The first video dataset is a Vimeo90K dataset, and the second video dataset is a LAVIB dataset.
8. The training method of the continuous motion based high-resolution video enhancement model according to claim 1, characterized in that: Before pre-training the video frame insertion model by adopting a training strategy with random frame insertion time and a preset multiple loss function, the method further includes: Performing data enhancement processing on the first video data set and setting training parameters; Before optimizing and training the pre-trained video frame insertion model, the method further includes: The data enhancement processing is performed on the second video data set, and training parameters are set.
9. The training method of the continuous motion based high-resolution video enhancement model according to claim 1, characterized in that: The multiple loss function is a weighted sum of a Laplace loss function and an optical flow loss function.
10. A video frame insertion method, characterized in that: include: Obtain the first frame image and the second frame image to be inserted; The first frame image, the second frame image and the interpolation time are input into a pre-trained video interpolation model, and an inserted intermediate frame image is output; wherein the pre-trained video interpolation model is obtained according to the training method of the high-resolution video enhancement model based on continuous motion according to any one of claims 1 to 9.
Citation Information
Patent Citations
Video frame insertion method and system based on full-to-multi-field transformation
CN116546237A
Multi-modal high-frame-rate frame insertion method based on edge enhancement
CN117097858A
Interpolation model learning method and device for learning interpolation frame generating module
US20240203113A1
Cited By
Road topology construction method and system based on interpolation timing enhancement
CN122510803A