Gradient guide optimization-based large-amplitude motion video frame insertion method and system

By adopting a network framework model that combines the network and gradient branch module across scales in video interpolation technology, the problem of difficult to determine the correspondence relationship of objects in large-scale motion video interpolation frames is solved, and a high-quality, natural and delicate interpolation effect is achieved.

CN120091158APending Publication Date: 2025-06-03SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510122367.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art has limitations when dealing with large-scale motion video interpolation, and it is difficult to adapt to small objects that move quickly, resulting in difficult to determine the correspondence between objects. The calculation complexity of traditional methods increases, making it difficult to achieve natural and delicate interpolation effects.

Method used

The network framework model consisting of video frame interpolation branch module and gradient branch module is adopted, and the cross-scale fusion network is used to fuse pictures of different scales to reduce error accumulation, and provide gradient guidance through the gradient branch module to optimize the generation of target frames.

Benefits of technology

The model's generation quality of video edge texture and local details is improved, and the subtle texture characteristics of the high-resolution layer are maintained, making the interpolated frames more natural and delicate, and suitable for large-scale motion scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091158A_ABST
    Figure CN120091158A_ABST
Patent Text Reader

Abstract

The invention discloses a gradient guide optimization-based large-amplitude motion video frame insertion method and system. The method comprises the following steps of: firstly, obtaining a large-amplitude motion video input frame; constructing a network framework model comprising a video frame interpolation branch module and a gradient branch module; and finally, executing a frame insertion task of a large-amplitude motion video input frame by utilizing the network framework model so as to obtain a gradient guide optimized target frame. A network framework model comprising a video frame interpolation branch module and a gradient branch module is utilized, a cross-scale fusion network is used for fusing pictures of different scales, error accumulation is reduced, the generation quality of the model for video edge textures and local details is enhanced, fine texture features of a high-resolution layer are kept, and the image quality is improved. And the frame insertion is more natural and fine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video frame interpolation, and particularly relates to a method and system for video frame interpolation of large-scale motion based on gradient-guided optimization. Background Art

[0002] Video Frame Interpolation aims to generate intermediate frames between given input frames, so as to convert low-frame-rate (LFR) content into high-frame-rate (HFR) video, in order to reduce motion blur while increasing the frame rate. This technology has important value in many practical applications, such as slow-motion video production, video compression, adaptive streaming transmission, frame rate improvement of video restoration, and new view synthesis. Currently, optical flow-based methods dominate the field of video frame interpolation, generating intermediate frames by estimating the optical flow from the target frame to the input frame and performing frame warping. However, due to the lack of a clear target frame, existing optical flow algorithms cannot be directly applied to the estimation of intermediate optical flow, which poses a challenge to complex scenes, especially scenes with large-scale motion.

[0003] The video frame interpolation task of large-scale motion faces more significant problems. In the above application scenarios, the pixel displacement between objects is large, resulting in difficult determination of the correspondence of objects between frames. Traditional methods often perform well in dealing with small-scale motion, but have limitations in large-scale motion scenarios. For example, the excessive depth of the model leads to an increase in computational complexity and difficulty in adapting to small objects with fast motion, which has become a technical problem that urgently needs to be solved. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for video frame interpolation of large-scale motion based on gradient-guided optimization to solve the deficiencies in the prior art. It uses a network framework model including a video frame interpolation branch module and a gradient branch module, and uses a cross-scale fusion network to fuse pictures of different scales, reduce error accumulation, enhance the generation quality of the model for video edge textures and local details, not only maintain the fine texture features of the high-resolution layer, but also make the frame interpolation more natural and delicate.

[0005] An embodiment of the present application provides a method for video frame interpolation of large-scale motion based on gradient-guided optimization, and the method includes:

[0006] Obtain the input frames of the large-scale motion video;

[0007] Construct a network framework model including a video frame interpolation branch module and a gradient branch module;

[0008] Use the network framework model to perform the frame interpolation task of the large-scale motion video input frames to obtain the target frames optimized by gradient guidance.

[0009] Optionally, the video frame interpolation branch module is implemented based on ST-MFNet, and the video frame interpolation branch module at least includes:

[0010] a multi-scale optical flow extraction module, a multi-scale fusion module, and a texture enhancement network that are communicatively connected; where

[0011] the multi-scale optical flow extraction module is used to generate multi-scale optical flow and features, and generate estimation results of different-scale front and back frames for the intermediate frame through warping;

[0012] the multi-scale fusion module is used to generate an intermediate result based on the GridNet architecture and the estimation result of the intermediate frame;

[0013] the texture enhancement network is used to output a residual signal containing the texture difference between the intermediate result and the target frame.

[0014] Optionally, the multi-scale optical flow extraction module includes:

[0015] a multi-scale intermediate optical flow estimation network and a bidirectional optical flow estimation network, where

[0016] the multi-scale intermediate optical flow estimation network is used to extract multi-scale multi-domain intermediate flows;

[0017] the bidirectional optical flow estimation network is used to extract large motion information.

[0018] Optionally, the multi-scale intermediate optical flow estimation network includes:

[0019] a feature extractor based on the U-Net style, the feature extractor includes eight MSResNext blocks, and two ResNext blocks are used in parallel for each MSResNext block. The kernel sizes of the intermediate layers of the ResNext blocks are 3×3 and 7×7 respectively.

[0020] Optionally, the extraction of the large motion information includes:

[0021] obtaining the input frames I 1 、I 2 corresponding to the bidirectional flows F 1→2 、F 2→1 ;

[0022] respectively obtaining approximate intermediate flows of the bidirectional flows F 1→2 、F 2→1 by using a preset linear method; where the preset linear method is: F 1→t = 0.5F 1→2 F 2→t = 0.5F 2→1 ;

[0023] Perform forward warping on the large - motion video input frame I according to the approximate intermediate flow using the softsplat operator 1 and I 2 to determine the large - motion information.

[0024] Optionally, performing the interpolation task of the large - motion video input frame using the network framework model to obtain the target frame optimized by gradient guidance, including:

[0025] Based on the network framework model and the large - motion video input frame I 1 and I 2 , obtain an intermediate result

[0026] Combine the intermediate result with the large - motion video input frame I 1 and I 2 in chronological order to generate a residual signal;

[0027] Based on the residual signal and the intermediate result determine the output of the video frame interpolation branch module

[0028] Use the gradient branch module to obtain the gradient branch result based on the gradient map and gradient features;

[0029] According to the output of the video frame interpolation branch module and integrate the gradient branch result into the video frame interpolation branch module to reversely guide the video frame interpolation branch module to generate the target frame optimized by gradient guidance.

[0030] Optionally, the gradient branch module is used to estimate the gradient map conversion of the target frame, and the output of the gradient branch module is used to calculate the following loss function to achieve the constraint on the target frame; where the loss function is:

[0031]

[0032] where M(·) represents the operation for extracting the gradient map, and ∈ = 0.001.

[0033] Another embodiment of the present application provides a large - motion video interpolation system based on gradient guidance optimization, and the system includes:

[0034] An acquisition module, configured to acquire large - motion video input frames;

[0035] A building module for building a network framework model including a video frame interpolation branch module and a gradient branch module;

[0036] An execution module for using the network framework model to perform an interpolation task on the input frames of a large - motion video to obtain a target frame optimized by gradient guidance.

[0037] Another embodiment of the present application provides a storage medium in which a computer program is stored. Wherein, the computer program is set to implement the method described in any one of the above when running.

[0038] Another embodiment of the present application provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is set to run the computer program to implement the method described in any one of the above.

[0039] Compared with the prior art, the present invention first obtains the input frames of a large - motion video; builds a network framework model including a video frame interpolation branch module and a gradient branch module; and finally uses the network framework model to perform an interpolation task on the input frames of a large - motion video to obtain a target frame optimized by gradient guidance. It uses a network framework model including a video frame interpolation branch module and a gradient branch module, and uses a cross - scale fusion network to fuse pictures of different scales, reducing error accumulation, enhancing the generation quality of the video edge texture and local details by the model. It not only preserves the fine texture features of the high - resolution layer, but also makes the interpolation more natural and delicate. Description of the Drawings

[0040] Figure 1 It is a hardware structure block diagram of a computer terminal for a large - motion video interpolation method based on gradient guidance optimization provided by an embodiment of the present invention;

[0041] Figure 2 It is a flow schematic diagram of a large - motion video interpolation method based on gradient guidance optimization provided by an embodiment of the present invention;

[0042] Figure 3 It is a structure schematic diagram of a network framework model provided by an embodiment of the present invention;

[0043] Figure 4 It is a structure schematic diagram of a video frame interpolation branch module provided by an embodiment of the present invention;

[0044] Figure 5 It is a structure schematic diagram of a multi - scale intermediate optical flow estimation network module provided by an embodiment of the present invention;

[0045] Figure 6Schematic diagram of a structure for connecting the output of a ResNext block to a channel attention module provided by an embodiment of the present invention;

[0046] Figure 7 Schematic diagram of a structure of a GridNet architecture module provided by an embodiment of the present invention;

[0047] Figure 8 Schematic diagram of a structure of a texture enhancement network architecture provided by an embodiment of the present invention;

[0048] Figure 9 Schematic diagram of a structure of a large - motion video frame interpolation system based on gradient - guided optimization provided by an embodiment of the present invention. Detailed implementation manners

[0049] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0050] An embodiment of the present invention first provides a large - motion video frame interpolation method based on gradient - guided optimization. This method can be applied to electronic devices, such as computer terminals, specifically, ordinary computers, tablets, etc.

[0051] The following takes running on a computer terminal as an example to explain it in detail. Figure 1 Hardware structure block diagram of a computer terminal for a large - motion video frame interpolation method based on gradient - guided optimization provided by an embodiment of the present invention. As Figure 1 shown, this computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non - volatile storage medium and an internal memory.

[0052] The non - volatile storage medium can store an operating system and a computer program. This computer program includes program instructions. When the program instructions are executed, the processor can execute any large - motion video frame interpolation method based on gradient - guided optimization.

[0053] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0054] The internal memory provides an environment for the operation of the computer program in the non - volatile storage medium. When this computer program is executed by the processor, the processor can execute any large - motion video frame interpolation method based on gradient - guided optimization.

[0055] This network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand, Figure 1The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0056] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0057] See Figure 2 , Figure 2 which is a schematic flowchart of a large motion video frame interpolation method based on gradient-guided optimization provided by an embodiment of the present invention, and may include the following steps:

[0058] S201: Obtain a large motion video input frame.

[0059] Specifically, to obtain a large motion video input frame, the input frame is a continuous video frame sequence with significant motion changes. Each video frame contains image information represented by a pixel matrix, and there may be large object movements, scene switches, or perspective changes between video frames.

[0060] S202: Construct a network framework model including a video frame interpolation branch module and a gradient branch module.

[0061] Specifically, see Figure 3 , Figure 3 which is a schematic structural diagram of a network framework model provided by an embodiment of the present invention, including a multi-scale optical flow extraction module, a multi-scale fusion module, a texture enhancement network (TENet), and two Warp modules that are communicatively connected. Figure 3 There are two branches in the overall network framework in , one of which is a video frame interpolation (VFI) branch module, and the other is a gradient branch module.

[0062] See Figure 4 , Figure 4Schematic diagram of the structure of a video frame interpolation branch module provided by an embodiment of the present invention. The video frame interpolation branch module is implemented based on ST-MFNet. The video frame interpolation branch module may at least include: a multi-scale optical flow extraction module, a multi-scale fusion module, and a texture enhancement network that are communicatively connected. Among them, the multi-scale optical flow extraction module is used to generate multi-scale optical flow and features, and generate estimation results of different-scale front and rear frames for the intermediate frame through warping. The multi-scale fusion module is used to generate an intermediate result based on the GridNet architecture and the estimation result of the intermediate frame. The texture enhancement network is used to output a residual signal including the texture difference between the intermediate result and the target frame.

[0063] Among them, the multi-scale optical flow extraction module includes: a multi-scale intermediate optical flow estimation network (Multi-InterFlow Network: MIFNet) and a bi-directional optical flow estimation network (Bi-directional Linear Flow Network: BLFNet). Among them, the multi-scale intermediate optical flow estimation network is used to extract multi-scale multi-domain intermediate flows. The bi-directional optical flow estimation network is used to extract large-motion information. It should be noted that the multi-scale intermediate optical flow estimation network may include: a feature extractor U-MultiScaleResNext (UMSResNext) based on the U-Net style. The feature extractor includes eight MSResNext blocks, and each MSResNext block uses two ResNext blocks in parallel. The intermediate layer kernel sizes of the ResNext blocks are 3×3 and 7×7 respectively.

[0064] Among them, the extraction of the large-motion information may include:

[0065] Step 1: Obtain the large-motion video input frames I 1 、I 2 corresponding bi-directional flows F 1→2 、F 2→1 ;

[0066] Step 2: Respectively obtain approximate intermediate flows of the bi-directional flows F 1→2 、F 2→1 by using a preset linear method. Among them, the preset linear method is: F 1→t = 0.5F 1→2 F 2→t = 0.5F 2→1 ;

[0067] Step 3: According to the approximate intermediate flow, use the softsplat operator to perform the large-motion video input frames I 1 、I 2Forward warping is performed to determine large motion information.

[0068] Exemplarily, for a given large motion video input frame I 1 、I 2 ,a network framework model including a video frame interpolation branch module and a gradient branch module first processes the large motion video input frames I 1 、I 2 in two branches. The multi-scale intermediate optical flow estimation network branch estimates the multi-scale multi-flows from I t to I 1 、I 2 ,where the many-to-one pixel correspondence allows complex transformations, which is beneficial for the interpolation of highly complex motions, such as dynamic textures (e.g., water, fire, etc.). Since the fixed receptive field of MIFNet may lead to limited ability to capture large motions, a bidirectional optical flow estimation network branch can also be added to approximately estimate the one-to-one optical flow from I 1 、I 2 to I t using a coarse-to-fine method to enhance the capture of large motions. The input frames are warped according to the flows generated by MIFNet and BLFNet, and then fused by the multi-scale fusion module to obtain an intermediate result This multi-branch structure combines the advantages of single-flow and multi-flow methods and is found to provide enhanced interpolation performance. In the second stage, is combined with all the input I 1 、I 2 in chronological order and input into the texture enhancement network (TENet), which captures longer-range dynamics and generates a residual signal for the final output.

[0069] Among them, the main role of MIFNet is to extract multi-scale multi-domain intermediate flows. Thus, a part mainly focuses on the extraction of detailed features and textures of the front and back frames for supporting the final multi-scale information fusion. The multi-interflow can be defined as G A→B =(α,β,ω), where α,β∈R H×W×N respectively represent the sets of the x and y components of N flow vectors, and ω∈[0,1] H×W×N is their weight That is to say, for each position (x,y), G A→B contains N flow vectors and N weights. The pixel value at each position (x,y) corresponding to the warping is defined as follows:

[0070]

[0071] See Figure 5 , Figure 5Schematic diagram of the structure of a multi-scale intermediate optical flow estimation network module provided by an embodiment of the present invention, where Figure 5 (a) is the overall architecture of the multi-scale intermediate optical flow estimation network module, which has a backbone with a U-Net structure and multi-flow estimation heads at three scales; Figure 5 (b) is the convolutional layer within each scale's multi-flow head. To capture multi-scale pixel motion, a U-Net style feature extractor U-MultiScaleResNext (UMSResNext) is designed, which consists of eight MSResNext blocks. Each MSResNext block uses two ResNext blocks in parallel, and the kernel sizes of the intermediate layers are different, being 3×3 and 7×7 respectively, which further increases the network cardinality.

[0072] See Figure 6 , Figure 6 Schematic diagram of the structure of connecting the output of the ResNext block to the channel attention module provided by an embodiment of the present invention. Here, the outputs of these two ResNext blocks are connected and then connected to the channel attention module, which learns to adaptively weight the feature maps extracted by the two ResNext blocks. This feature selection mechanism is also found to enhance motion modeling. In UMSResNext, the upsampling operation is performed by replacing the k×k grouped convolution of the intermediate layer with a (k + 1)×(k + 1) grouped transposed convolution. Then, the features extracted by UMSResNext are passed to the multi-flow head for inter-multi-flow prediction. Here, the multi-flow is predicted at multiple scales l = -1, 0, 1. As Figure 5 (b) shows, each multi-flow head contains 6 sub-branches, predicting the x, y components (α, β) and kernel weights (ω) of G t→1 , G t→2 . Then, the predicted flow is used to perform inverse warping on the inputs I 1 , I 2 at the corresponding scale.

[0073] The main function of BLFNet is to extract large motion information. The approach is to predict the bidirectional flows F 1 , F 2 between the inputs I 1→2 , F 2→1 , and then linearly approximate the intermediate flow in the following way. According to the intermediate flow, the frames I 1 , I 2 are forward warped using the efficient softsplat operator:

[0074] F 1→t = 0.5F 1→2 F 2→t = 0.5F 2→1

[0075] In summary, in the previous large - motion video frame interpolation techniques, the optical flow was often optimized in a multi - scale progressive manner, which led to the accumulation of errors. At the same time, it only focused on the refinement of the optical flow and ignored the use of image texture information. This application uses a cross - scale fusion network to fuse images at different scales, which has the following advantages compared with the traditional use of pyramid optical flow refinement:

[0076] 1. Reducing error accumulation: In the traditional pyramid optical flow, the information is passed layer by layer from coarse to fine. If the error in the low - resolution layer is large, it is difficult to completely correct it in the subsequent layers. However, the cross - scale fusion network can correct the possible early deviations at each scale by fusing the features of the same frame at different resolutions in parallel or interactively, significantly reducing the error cascading in recursive refinement.

[0077] 2. Better retaining texture details: The pyramid method weakens the high - frequency information in the low - resolution stage and it is difficult to fully compensate for it later. The cross - scale fusion network processes in parallel at multiple resolution levels, can capture both local fine textures and overall large - motion information simultaneously, retains the subtle texture features of the high - resolution layer and flexibly adjusts the estimation results, making the frame interpolation more natural and delicate.

[0078] It should be noted that the multi - scale fusion module is used to generate intermediate interpolation results using the frames distorted at multiple scales in the previous steps. Here, this application can adopt the GridNet architecture because it performs excellently in fusing multi - scale information. See Figure 7 , Figure 7 which is a schematic structural diagram of a GridNet architecture module provided by an embodiment of the present invention. Here, the GridNet is configured to have 4 columns and 3 rows. The first, second, and third rows correspond to scales of l = - 1, 0, 1 respectively. The first row and the third row take and as inputs, while the second row takes as the input, where {·} represents channel - by - channel connection. Finally, this module outputs the intermediate result at the original spatial resolution l = 0.

[0079] The output of the multi - scale fusion module is connected to the original input and the guiding information provided by the gradient branch, and then they are input into the texture enhancement network. Here, including additional frames can better model high - order motion and can also provide more information about long - term spatio - temporal features. Inspired by recent research in the field of dynamic texture synthesis, it is found that spatio - temporal filtering is very effective for generating coherent video textures, and a 3D CNN is integrated to enhance the texture. See Figure 8 , Figure 8It is a schematic structural diagram of a texture enhancement network architecture provided by an embodiment of the present invention. The texture enhancement network needs to output a residual signal containing the texture difference between and the target frame. Adding the residual to is the final output of the video frame interpolation branch module

[0080] S203: Using the network framework model, perform the frame interpolation task on the input frames of the large motion video to obtain a target frame optimized by gradient guidance.

[0081] Specifically, the using the network framework model to perform the frame interpolation task on the input frames of the large motion video to obtain a target frame optimized by gradient guidance may include:

[0082] 1. Based on the network framework model and the input frames I 1 、I 2 of the large motion video, obtain an intermediate result

[0083] 2. Combine the intermediate result with the input frames I 1 、I 2 of the large motion video in chronological order to generate a residual signal;

[0084] 3. Based on the residual signal and the intermediate result determine the output of the video frame interpolation branch module

[0085] 4. Using the gradient branch module, obtain a gradient branch result based on the gradient map and gradient features;

[0086] 5. According to the output of the video frame interpolation branch module and integrate the gradient branch result into the video frame interpolation branch module to reversely guide the video frame interpolation branch module to generate a target frame optimized by gradient guidance.

[0087] Specifically, when dealing with the task of large - motion video frame interpolation, the foreground occlusion of the background caused by large motion amplitudes often leads to problems such as inaccurate foreground - background segmentation and incomplete background restoration. Gradient guidance can provide significant advantages here: on the one hand, by emphasizing the gradient mutation regions, it highlights the object edges and structural information, enabling the model to more accurately identify the foreground object contours and effectively segment them from the background; on the other hand, adding constraints in the gradient dimension during the deep - network training can provide a more explicit supervision signal for occlusion repair and background restoration, thereby enhancing the model's generation quality of edge textures and local details; in addition, by integrating gradient guidance into a multi - stage or multi - branch frame - interpolation pipeline, it can continuously calibrate the model's perception and reconstruction of the foreground occlusion region during the interaction between motion estimation and occlusion repair, making the frame - interpolation results still maintain a smooth, natural, and sharp - edged visual effect in large - range motion scenarios. Therefore, a gradient - branch module is set up in this application to guide and optimize the interpolation results.

[0088] As Figure 3 shown, the gradient - branch module combines multiple intermediate representation layers from the video - frame interpolation - branch module. The motivation for this scheme is that the carefully designed video - frame interpolation - branch module can carry rich structural information, which is crucial for the recovery of the gradient map. Therefore, these features are used as powerful prior knowledge to improve the performance of the gradient branch while significantly reducing the parameters of the gradient branch in this case. When generating the intermediate - frame gradient map in the Warp module, the optical flow generated by the video - frame interpolation - branch module can be used because when generating optical flow using the gradient map, due to the lack of relevant information in the gradient map, the optical flow generated using the gradient map has a large error, and at the same time, this can reduce the complexity of the model. Once the gradient map is obtained through the gradient - branch module, the obtained gradient features can be integrated into the video - frame interpolation - branch module to inversely guide the VFI to generate the target frame. The magnitude of the gradient map can implicitly reflect whether the restored region should be clear or smooth. In practical applications, the feature map generated by the penultimate layer of the gradient branch can be input into the video - frame interpolation - branch module. The gradient - branch module finally outputs for calculating the following loss function to impose constraints on the final result.

[0089] Exemplarily, the gradient - branch module is used to estimate the gradient - map transformation of the target frame, and the output of the gradient - branch module is used to calculate the following loss function to achieve constraints on the target frame; where the loss function can be expressed as:

[0090]

[0091] where M(·) represents the operation for extracting the gradient map, and ∈ = 0.001.

[0092] It should be noted that the goal of the gradient branch module is to estimate the gradient map transformation of the target frame, and the gradient map can be obtained by calculating the differences between adjacent pixels:

[0093] I x (x) = I(x + 1, y) - I(x - 1, y)

[0094] I y (x) = I(x, y + 1) - I(x, y - 1)

[0095]

[0096] Among them, the elements of the gradient map are the gradient lengths of pixel coordinates x = (x, y), and the gradient can be easily obtained through a convolutional layer with a fixed convolution kernel. In fact, the vector direction information of the gradient can be ignored because the gradient intensity (absolute value) is sufficient to reveal the sharpness of local regions in the restored image. Therefore, the intensity map is used as the gradient map. Such a gradient map can be regarded as another kind of image, so the image-to-image conversion technology can be used to learn the mapping between the two modalities. Since most regions of the gradient map are close to zero, the convolutional neural network can focus more on the spatial relationships of the contours. Therefore, the network may be more likely to capture the structural dependencies, thereby generating an approximate gradient map of the target frame.

[0097] It can be seen that the present invention first obtains the input frames of the large-motion video; constructs a network framework model including a video frame interpolation branch module and a gradient branch module; and finally uses the network framework model to perform the frame interpolation task on the input frames of the large-motion video to obtain a target frame optimized by gradient guidance. It uses a network framework model including a video frame interpolation branch module and a gradient branch module, and uses a cross-scale fusion network to fuse pictures of different scales, reducing error accumulation, enhancing the generation quality of the video edge texture and local details of the model, not only maintaining the fine texture features of the high-resolution layer, but also making the frame interpolation more natural and delicate.

[0098] Furthermore, the network framework model including the video frame interpolation branch module and the gradient branch module in this application is compared with the currently most advanced VFI models in each field on the general task video dataset and the large-motion video dataset respectively. The quantitative comparison results are shown in Tables 1 and 2. Table 1 shows the comparison results of the network framework model of this application with relevant VFI models on the general video dataset, and Table 2 shows the comparison results of the network framework model of this application with relevant VFI models on the large-motion video dataset. For fair comparison, all benchmark models were retrained using the same training and validation datasets as ST-MFNet under the same training configuration.

[0099] Table 1: Comparison Results Table of This Application with Related VFI Models on General Task Video Datasets (PSNR / SSIM)

[0100]

[0101] Table 2: Comparison Results Table of This Application with Related VFI Models on Large Motion Video Datasets (PSNR / SSIM)

[0102]

[0103]

[0104] Among them, "OOM" indicates that an out-of-memory problem occurred during evaluation on an NVIDIA 4090-24G GPU.

[0105] Two key observations can be drawn from Table 1 and Table 2. First, by using the preset training set (Vimeo-90k + BVI-DVC), on the general task video dataset, the network framework model of this application provides the best results on SNU-FILM (all subsets). Compared with the runner-up of each test set, the PSNR has increased significantly by 0.47 - 1.078 dB. Its performance on DAVIS is only better than the pre-trained FLAVR, with a marginal difference of 0.177 dB (PSNR). Moreover, the network framework model of this application is higher than the existing two large motion video interpolation models on the large motion video dataset. Compared with the runner-up of each test set, the PSNR has increased significantly by 0.470 - 1.064 dB. This proves that the proposed ST-MFNet has excellent generalization ability.

[0106] Therefore, for the large motion video interpolation framework described in this application, when dealing with the large motion video interpolation task, due to the large motion amplitude, the foreground occlusion of the background often causes problems such as inaccurate foreground-background segmentation and incomplete background restoration. Gradient guidance can provide significant advantages here: on the one hand, by emphasizing the gradient mutation region, it highlights the object edges and structural information, enabling the model to more accurately identify the foreground object contour and effectively segment it from the background; on the other hand, adding constraints in the gradient dimension during the deep network training can provide a more explicit supervision signal for occlusion repair and background restoration, thereby enhancing the model's generation quality of edge textures and local details. Therefore, a gradient branch is set in this method to guide and optimize the interpolation results.

[0107] In the use of a cross-scale fusion module in large motion video frame interpolation, in traditional pyramid optical flow, information is passed layer by layer from coarse to fine. If the error in the low-resolution layer is large, it is difficult to fully correct it subsequently. In contrast, the cross-scale fusion network can fuse features of the same frame at different resolutions in parallel or interactively, and can correct possible early biases at each scale, significantly reducing the error cascades that occur in recursive refinement. The pyramid method weakens high-frequency information in the low-resolution stage and it is difficult to fully compensate for it subsequently. The cross-scale fusion network processes in parallel at multiple resolution levels, can capture both local fine textures and overall large motion information simultaneously, preserves the fine texture features of the high-resolution layer, and flexibly adjusts the estimation results, making the frame interpolation more natural and delicate.

[0108] Motion compensation is achieved for gradient image prediction using the optical flow predicted from the original image. When predicting the intermediate frame in the gradient branch, since there is a certain loss of texture information in the gradient map compared to the original image, there will be a large error when using the gradient map to extract the optical flow, resulting in a reduction in the quality of the final interpolation result. Therefore, in the method described in this application, the optical flow extracted from the original image is used to perform a deformation operation on the gradient, thereby achieving motion compensation for the deformed gradient image. Another advantage of this approach is that it can reduce the complexity of the algorithm.

[0109] The large - motion video frame interpolation technology mainly generates intermediate frames between two video frames with large motion or scene changes through optical flow estimation or deep - learning models. It can not only improve the smoothness and visualization effect of videos, but also show significant application prospects in fields such as film and television production, game rendering, sports analysis, industrial inspection, medical imaging, and communication. For the film and television industry, it can repair footage with insufficient shooting frame rates or old materials into higher - frame - rate images, enhancing the visual experience and reducing the cost of reshooting, and providing more possibilities for special - effects production and historical data preservation. In the game field, frame interpolation technology can help cloud games or VR / AR displays achieve smooth high - frame - rate images, thus bringing a more immersive experience. At the same time, sports broadcasts and replays can also achieve clear slow - motion and precise action - detail analysis through this technology, making the work of commentators and referees more efficient. In the industrial aspect, on the production line, frame interpolation can make up for the deficiencies of low - frame - rate monitoring or inspection videos and capture subtle defects in a timely manner. In scientific research experiments, most high - speed motion recordings and analyses can be completed with ordinary video equipment and frame - interpolation algorithms without expensive high - speed cameras. In medical video training and auxiliary diagnosis, frame interpolation can also enhance the visualization of surgical procedures and the observation of dynamic lesions, providing better support for teaching and clinical practice. In communication and social platforms, frame interpolation can improve the smoothness of video calls or short videos in scenarios with limited bandwidth, allowing users to enjoy a better viewing and interaction experience. More forward - looking applications also include emerging fields such as autonomous driving, 3D vision, multi - view synthesis, and content - adaptive compression, supplementing richer motion information for vehicle sensors or panoramic videos and dynamically reconstructing high - frame - rate images on the terminal side. With the continuous evolution of deep - learning and hardware - acceleration technologies, large - motion video frame interpolation will continue to expand new application scenarios and become a key driving force for enriching content creation, improving image quality, and enhancing the user experience.

[0110] Compared with the prior art, the present invention first obtains the input frames of large - motion videos; constructs a network - framework model including a video - frame interpolation branch module and a gradient branch module; and finally uses the network - framework model to perform the frame - interpolation task on the input frames of large - motion videos to obtain the target frames optimized by gradient guidance. It uses a network - framework model including a video - frame interpolation branch module and a gradient branch module, and uses a cross - scale fusion network to fuse pictures of different scales, reducing error accumulation and enhancing the generation quality of the model for video edge textures and local details. It not only preserves the fine texture features of the high - resolution layer, but also makes the frame interpolation more natural and delicate.

[0111] Another embodiment of the present application provides a large - motion video frame interpolation system based on gradient - guidance optimization, as Figure 9 shown in the structural schematic diagram of a large - motion video frame interpolation system based on gradient - guidance optimization. The system includes:

[0112] An acquisition module 901, configured to acquire input frames of a large-scale motion video;

[0113] A construction module 902, configured to construct a network framework model including a video frame interpolation branch module and a gradient branch module;

[0114] An execution module 903, configured to use the network framework model to perform an interpolation task on the input frames of the large-scale motion video to obtain target frames optimized by gradient guidance.

[0115] Compared with the prior art, the present invention first acquires input frames of a large-scale motion video; constructs a network framework model including a video frame interpolation branch module and a gradient branch module; and finally uses the network framework model to perform an interpolation task on the input frames of the large-scale motion video to obtain target frames optimized by gradient guidance. It uses a network framework model including a video frame interpolation branch module and a gradient branch module, and uses a cross-scale fusion network to fuse pictures of different scales, reduces error accumulation, enhances the generation quality of the video edge texture and local details of the model, not only maintains the fine texture features of the high-resolution layer, but also makes the interpolation more natural and delicate.

[0116] An embodiment of the present invention further provides a storage medium, in which a computer program is stored, and the computer program is configured to implement the steps in the above method embodiment when running.

[0117] Specifically, in this embodiment, the above storage medium may be configured to store a computer program for performing the following steps:

[0118] S201: Acquire input frames of a large-scale motion video;

[0119] S202: Construct a network framework model including a video frame interpolation branch module and a gradient branch module;

[0120] S203: Use the network framework model to perform an interpolation task on the input frames of the large-scale motion video to obtain target frames optimized by gradient guidance.

[0121] Specifically, in this embodiment, the above storage medium may include but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0122] Compared with the prior art, the present invention first obtains a large - motion video input frame; constructs a network framework model including a video frame interpolation branch module and a gradient branch module; and finally uses the network framework model to perform an interpolation task on the large - motion video input frame to obtain a target frame optimized by gradient guidance. It uses a network framework model including a video frame interpolation branch module and a gradient branch module, and uses a cross - scale fusion network to fuse pictures of different scales, reducing error accumulation, enhancing the generation quality of the video edge texture and local details by the model, not only maintaining the fine texture features of the high - resolution layer, but also making the interpolation more natural and delicate.

[0123] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in the above - mentioned method embodiment.

[0124] Specifically, the above - mentioned electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above - mentioned processor, and the input / output device is connected to the above - mentioned processor.

[0125] Specifically, in this embodiment, the above - mentioned processor may be configured to execute the following steps through a computer program:

[0126] S201: Obtain a large - motion video input frame;

[0127] S202: Construct a network framework model including a video frame interpolation branch module and a gradient branch module;

[0128] S203: Use the network framework model to perform an interpolation task on the large - motion video input frame to obtain a target frame optimized by gradient guidance.

[0129] Compared with the prior art, the present invention first obtains a large - motion video input frame; constructs a network framework model including a video frame interpolation branch module and a gradient branch module; and finally uses the network framework model to perform an interpolation task on the large - motion video input frame to obtain a target frame optimized by gradient guidance. It uses a network framework model including a video frame interpolation branch module and a gradient branch module, and uses a cross - scale fusion network to fuse pictures of different scales, reducing error accumulation, enhancing the generation quality of the video edge texture and local details by the model, not only maintaining the fine texture features of the high - resolution layer, but also making the interpolation more natural and delicate.

[0130] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0131] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0132] In the several embodiments provided by the present invention, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0133] The units described as separate components above may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0134] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0135] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of the present invention. The aforementioned memory includes various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0136] The embodiments of the present invention have been described in detail above. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for interpolating frames in large-scale motion video based on gradient-guided optimization, characterized in that: The method comprises: Obtaining large motion video input frames; Construct a network framework model including a video frame interpolation branch module and a gradient branch module; The network framework model is used to perform the interpolation task of large-scale motion video input frames to obtain gradient-guided optimized target frames.

2. The method according to claim 1, characterized in that The video frame interpolation branch module is implemented based on ST-MFNet, and the video frame interpolation branch module at least includes: A multi-scale optical flow extraction module, a multi-scale fusion module and a texture enhancement network connected by communication; wherein, The multi-scale optical flow extraction module is used to generate multi-scale optical flows and features, and generate estimation results of the front and back frames of different scales for the intermediate frame by distortion; The multi-scale fusion module is used to generate an intermediate result based on the GridNet architecture and the estimation result of the intermediate frame; The texture enhancement network is used to output a residual signal containing the texture difference between the intermediate result and the target frame.

3. The method according to claim 2, characterized in that The multi-scale optical flow extraction module includes: Multi-scale intermediate optical flow estimation network and bidirectional optical flow estimation network, where: The multi-scale intermediate optical flow estimation network is used to extract multi-scale multi-domain intermediate flows; The bidirectional optical flow estimation network is used to perform extraction of large-scale motion information.

4. The method according to claim 3, characterized in that The multi-scale intermediate optical flow estimation network includes: Based on a U-Net style feature extractor, the feature extractor includes eight MSResNext blocks, and each MSResNext block uses two ResNext blocks in parallel, and the intermediate layer kernel sizes of the ResNext blocks are 3×3 and 7×7 respectively.

5. The method according to claim 4, characterized in that The step of extracting large-scale motion information comprises: Obtain the bidirectional flow F corresponding to the large-scale motion video input frames I1 and I2 1→2 、F 2→1 ; The bidirectional flow F is obtained by using a preset linear method. 1→2 、F 2→1 The approximate intermediate flow of ; wherein the preset linear method is: F 1→t =0.5F 1→2 F 2→t =0.5F 2→1 ; According to the approximate intermediate stream, forward warping of the large motion video input frames I1, I2 is performed using a softsplat operator to determine large motion information.

6. The method according to claim 5, characterized in that The method of using the network framework model to perform a frame insertion task of a video input frame with large motion to obtain a target frame for gradient-guided optimization includes: Based on the network framework model and the large-scale motion video input frames I1 and I2, an intermediate result is obtained. The intermediate results Combine with the large-motion video input frames I1 and I2 in time sequence to generate a residual signal; Based on the residual signal and the intermediate result Determine the output of the video frame interpolation branch module Using the gradient branching module, a gradient branching result based on a gradient map and gradient features is obtained; According to the output of the video frame interpolation branch module The gradient branch result is integrated into the video frame interpolation branch module to reversely guide the video frame interpolation branch module to generate a gradient-guided optimized target frame.

7. The method according to claim 6, characterized in that The gradient branch module is used to estimate the gradient map conversion of the target frame, and the output of the gradient branch module Used to calculate the following loss function to implement the constraints on the target frame; where the loss function is: Here, M(·) represents the operation for extracting the gradient map, ∈=0.

001.

8. A large-motion video interpolation system based on gradient-guided optimization, characterized in that: The system comprises: An acquisition module, used for acquiring a large-scale motion video input frame; A construction module, used to construct a network framework model including a video frame interpolation branch module and a gradient branch module; The execution module is used to use the network framework model to perform the interpolation task of the large-scale motion video input frame to obtain the target frame guided by gradient optimization.

9. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to implement the method according to any one of claims 1 to 7 when executed.

10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the computer program to implement the method according to any one of claims 1 to 7.