A method and system for motion refinement for video interpolation

By processing two frames of images using a motion refinement model, the occlusion and blurring problems in video frame interpolation methods for large motion scenes are solved, achieving high-quality video frame interpolation effects.

CN117058196BActive Publication Date: 2025-12-05NANHAI RES STATION OF INST OF ACOUSTICS CHINESE ACADEMY OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310923648.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-26
Publication Date
2025-12-05
Estimated Expiration
2043-07-26

AI Technical Summary

Technical Problem

Existing stream-based video frame interpolation methods suffer from occlusion and blurring issues when generating moving objects in high-motion scenes, and also have many model parameters and high computational complexity.

Method used

A motion thinning method is adopted, which processes two frames of images through a motion thinning model, including a downsampling module, a context feature extraction module, a joint stream coding module, and a motion-guided feature fusion module, to generate improved optical flow estimation and pixel reliability scores. The intermediate frame is then synthesized by combining pixel warping and fusion strategies.

Benefits of technology

It effectively handles occlusion and pixel blurring issues in high-motion scenes, improves the fidelity of moving objects and inter-frame consistency, and achieves high-quality video frame interpolation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058196B_ABST
    Figure CN117058196B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of machine vision and deep learning, and particularly relates to a motion refinement method with local compensation for video frame interpolation. The method comprises: on the basis of the original motion optical flow estimation improvement framework M2M-PWC, two modules are newly added, which are a context feature extraction module and a motion guided feature fusion module. The CCFE module is integrated into each layer of the pyramid structure, which aims to encourage the model to extract clean and sufficient context information from the input image. The MGFF can guide the image feature fusion based on the optical flow features, so that the feature fusion of the moving object is more accurate, thereby providing local compensation for the optical flow estimation. Through the above two modules, the enhanced motion optical flow estimation can be obtained, which describes the inter-frame motion field information in detail, can effectively handle the occlusion and pixel blur problems in the large motion scene, and realizes high-quality video frame interpolation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of machine vision and deep learning technology, and in particular to a method and system for motion refinement of video frame interpolation. Background Technology

[0002] Video frame interpolation, an application of computer vision in video enhancement, has attracted significant attention from scholars in recent years. The purpose of video frame interpolation is to increase the frame rate of a video by inserting one or more intermediate frames between adjacent frames in the original video sequence. Besides converting videos to higher frame rates and improving visual effects, video frame interpolation techniques can also be used for video compression, video editing, generating training data to learn how to synthesize motion estimation, and as an auxiliary task for optical flow estimation.

[0003] Existing research on video frame interpolation typically relies on deep neural networks and can be categorized into stream-based and kernel-based methods. Kernel-based video frame interpolation methods synthesize the target frame by predicting the interpolation kernel for each pixel, while stream-based methods synthesize the target frame by estimating optical flow to warp the frame. While kernel-based methods are efficient, they are limited to interpolating frames within a fixed time step, and their runtime increases linearly with the desired number of output frames. Motion-based video frame interpolation methods, by establishing dense correspondences between frames and applying warping to render intermediate pixels, can effectively reduce interpolation time and allow for arbitrary-time interpolation. Therefore, stream-based video frame interpolation methods have become the dominant approach for arbitrary-time interpolation.

[0004] Existing flow-based methods have achieved promising results in generating realistic and consistent inter-frame representations. However, these methods often employ increasingly complex networks, leading to a greater number of model parameters and higher computational complexity. The proposed M2M-PWC is a video interpolation method based on optical flow estimation, enabling arbitrary-time interpolation while significantly reducing model parameters and improving inference speed. However, M2M-PWC still has room for improvement; for example, in high-motion scenes, generated moving objects suffer from occlusion and blurring issues. Therefore, how to further improve the inter-frame optical flow estimation algorithm for large-motion scenes while maintaining the required model parameters remains a problem that needs further exploration. Summary of the Invention

[0005] The purpose of this invention is to overcome the technical defects of existing video frame interpolation methods and to propose a motion refinement method with local compensation for video frame interpolation.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solution.

[0007] This invention proposes a motion thinning method for video frame interpolation, the method comprising:

[0008] The two frames to be interpolated are processed to obtain the original optical flow estimate between the two frames, which is used as the optical flow feature;

[0009] Two frames of images and the original optical flow estimate are input into a pre-established and trained motion thinning model to achieve motion thinning, resulting in an improved optical flow estimate and pixel reliability score. Based on the improved optical flow estimate and pixel reliability score, an intermediate frame is synthesized through pixel warping and fusion strategies.

[0010] The motion refinement model includes: a downsampling module, a context feature extraction module, a joint stream coding module, a motion-guided feature fusion module, and a decoder; wherein...

[0011] The downsampling module uses two image pyramid structures to downsample the features of two frames of images respectively.

[0012] The context feature extraction module is used to enhance the image features of each layer of the downsampling module;

[0013] The joint stream coding module is used to interactively fuse image features and optical flow features;

[0014] The motion-guided feature fusion module is used to fuse pixel information based on prior knowledge of optical flow estimation and to process the features fused by the joint stream coding module to achieve fine fusion of guided optical flow features and image features.

[0015] The decoder is used to decode optical flow features and image features obtained from the joint stream coding module, context feature extraction module and feature fusion module, and generate an improved optical flow estimate and its corresponding pixel reliability score.

[0016] As one improvement to the above technical solution, each downsampling module is implemented using two two-dimensional convolutions with different parameters and a PReLu activation function; the image features I of the l-th layer downsampling module sl The expression is:

[0017] I sl =PReLu(Conv2d(PReLu(Cond2d(I s(l-1) ))))

[0018] Where s∈{0,1} represents the frame source, l∈{1,…,L} represents the first layer, L represents the total number of layers, and I s(l-1) The image features are represented by the (l-1)th downsampling layer, where Cond2d() represents a two-dimensional convolution and PReLu() represents the activation function.

[0019] As an improvement to the above technical solution, the context feature extraction module includes: a multi-scale feature aggregation mechanism block and a channel attention mechanism block, and the processing procedure specifically includes:

[0020] Step A1. Convert the feature map I obtained by the downsampling module sl The input is a multi-scale feature aggregation mechanism block, which consists of k dilated convolutional branches with different dilation rates. Each of the k dilated convolutional branches respectively processes image feature I. sl The process involves generating new feature maps, and the steps are as follows:

[0021] O k =dilated conv k (I sl )

[0022] Among them, dilated conv k () and O k Let represent the dilated convolution operation and output of the k-th branch, respectively;

[0023] Step A2. Concatenate the features generated by all branches along the channel dimension and pass them through the δ activation function to obtain a multi-scale feature map u. sl The expression is:

[0024] u sl =δ(Concat(O) k ))

[0025] Where δ() represents the ReLU activation function, and Concat represents concatenation;

[0026] Step A3. Put u sl The input channel attention mechanism block includes a global average pooling layer, two convolutional layers, and a multiplication operation layer.

[0027] The global average pooling layer performs a global average pooling operation on the feature map u. sl Compressed along the spatial dimension into a feature vector z csl The process expression is:

[0028]

[0029] Where G() represents the global average pooling operation, H×W represents the size of the input feature map, W represents the width, and H represents the height; u sl (x, y) represents the feature map u sl The eigenvalue at a height of x and a width of y;

[0030] Step A4. Transfer feature z cslThe vector is transformed into learnable parameters, and the specific formula is as follows:

[0031] v csl =T(z) csl ,w)=σ(Conv2(δ(Conv1(z csl )))

[0032] Where T(·) represents the vector transformation function, δ represents the ReLU activation function, and σ represents the sigmoid activation function; Conv1 and Conv2 are respectively the transformation functions of the feature vector z. csl One-dimensional and two-dimensional convolutions are used for mapping;

[0033] Step A4. The product operation layer weights the learned feature channel importance coefficients w with their corresponding feature channels to obtain the recalibrated features x. sl The formula is:

[0034] x sl =v csl ·u csl ,where c∈(0,1,…,C)

[0035] Where c and C represent the current number of channels and the total number of channels, respectively, v csl u represents the channel weight calculated at channel c in the l-th layer of image s. csl This represents the feature value of image s at channel c in the l-th layer.

[0036] As one improvement to the above technical solution, the processing procedure of the joint stream coding module specifically includes:

[0037] Step B1. From layer 1 to layer L, apply the joint stream coding module to generate the optical flow feature pyramid of the bidirectional flow field;

[0038] In the l-th layer of the joint stream coding module, the optical flow features and image features from the previous layer are distorted together. Specifically, the optical flow estimation from the previous layer is used, along with the corresponding pyramid features I. 0(l-1) To I 1(l-1) Distortion, obtaining the distortion feature wim 1l ; and the corresponding pyramid feature I 1(l-1) To I 0(l-1) Distortion, obtaining the distortion feature wim 0l The expressions are as follows:

[0039] wim 1l =warp(I 1(l-1) f 0(l-1) )

[0040] wim 0l =warp(I 0(l-1)f 1(l-1) )

[0041] Among them, warp() means warp;

[0042] Step B2. Concatenate and combine the original optical flow features, image features, and distortion features, and then downsample them using two convolutional layers:

[0043] e 0l =Down(Concat(f 0(l-1) I 0(l-1) wim 1l ))

[0044] e 1l =Down(Concat(f 1(l-1) I 1(l-1) wim 0l ))

[0045] Concat() represents concatenation, and Down() represents downsampling;

[0046] Step B3. Downsample the optical flow characteristics of layer l (l-1) to obtain the optical flow characteristics f of layer l on both branches. 0l and f 1l The expressions are as follows:

[0047] f 0l =Down(f 0(l-1) )

[0048] f 1l =Down(f 1(l-1) ).

[0049] As an improvement to the above technical solution, the processing procedure of the motion-guided feature fusion module specifically includes:

[0050] Step C1. The feature fusion module uses the hybrid feature e output from the Lth layer of the joint stream coding module. L Optical flow characteristics f L Image downsampling features I L After fusion, three multi-scale weighted features w1, w2, and w3 are obtained, expressed as follows:

[0051]

[0052]

[0053]

[0054] in, and Indicates a fully connected layer operation, Downp () represents the downsampling function implemented by a 2D convolution with a stride of 2, where p1 and p2 represent the downsampling ratios;

[0055] Step C2. Merge the weighted features to obtain p1 times upsampling weight w4 and p2 times upsampling weight w o :

[0056] w4 = w2 + Up d (w3)

[0057]

[0058] Among them, UP a () denotes the upsampling function implemented by bilinear interpolation, where d is the sampling factor; For a fully connected layer, σ() is the Sigmoid activation function;

[0059] Step C3. Obtain the fused feature e using the following formula. l+1 :

[0060] e L =e L ·w o +I L ·(1-w o ).

[0061] As one improvement to the above technical solution, the encoder's processing procedure specifically includes:

[0062] Step D1. The decoder obtains feature maps of the optical flow feature pyramid and the image feature pyramid based on the joint stream coding module and the feature fusion module, and generates N motion sub-vectors step by step using deconvolution. That is, the improved optical flow estimation and its corresponding pixel reliability score {S0, S1}, where m0, m1, m2, and m3 represent hidden layer features, and the calculation formula is:

[0063] m0 = PReLu(DeConv(Concat[e L c L f L ]))

[0064] m1 = PReLu(DeConv(Concat[e L-1 ,m0]))

[0065] m2=PReLu(DeConv(Concat[e L-2 ,m1]))

[0066] m3=PReLu(DeConv(Concat[e L-3 ,m2]))

[0067]

[0068] Where PReLu() represents the activation function and DeConv() represents deconvolution;

[0069] As one improvement to the above technical solution, the process of synthesizing intermediate frames based on improved optical flow estimation and pixel reliability scores through pixel warping and fusion strategies includes the following steps:

[0070] Step D2. First, combine the multiple motion vectors Scaling to a given target time step t:

[0071]

[0072]

[0073] Where i0 and i1 represent the i-th source pixel of I0 and I1, respectively; This represents the optical flow estimate of the nth motion vector at pixel i0 from source frame 0 to source frame l. This represents the optical flow estimate of the nth motion vector at pixel i1 from source frame 1 to source frame 0. Indicates will Optical flow estimate scaled to target time t Indicates will Optical flow estimate scaled to target time t;

[0074] Step D3. Then, a source pixel i s The nth motion vector will be warped forward by the forward warping operation φ to the source pixel i s Warped The expression is:

[0075]

[0076] Step D4. The nth motion sub-vector can warp each pixel in the source frame s forward to the target time t, thus obtaining... The set of target pixels after warping of N sub-motion vectors is:

[0077]

[0078] Step D5. Measure the importance of each pixel from three aspects: temporal correlation, brightness consistency, and reliability score:

[0079] Time correlation is represented by r i Let r represent the value of i when i is a pixel from I0. i = 1-t, and when i is a pixel from I1, ri =t;

[0080] Brightness uniformity is represented by b i It is indicated that it is obtained in the following way:

[0081]

[0082] The reliability scores {S0, S1} are obtained through step D1;

[0083] Step D6. Using pixel warping and fusion strategies, the original pixel i is fused to obtain the intermediate frame pixel j:

[0084]

[0085] Among them, s i It is the pixel reliability score at position i; c i α is the original color of pixel i; α is a learnable parameter used to adjust the importance of the weights.

[0086] As an improvement to the above technical solution, the method further includes: training the motion refinement model to obtain a trained motion refinement model; the training process specifically includes:

[0087] Step 1) Obtain data: Randomly read a set of triplet data. Each triplet contains three consecutive frames of images. The first and third frames are the input to the motion refinement model, and the second frame is the real label for supervised training.

[0088] Step 2) Obtain the original optical flow estimate. Input the first and third frames I0 and I1 from the triplet image into the existing optical flow estimation model PWC-NET, and estimate the original optical flow F in two directions between the two input frames. 0->1 and F 1->0 ;

[0089] Step 3) Take the first and third frames I0 and I1 from the triplet image, and the original optical flow estimate F 0->1 and F 1->0 The improved optical flow estimate and pixel reliability score are obtained by inputting the data into the established motion refinement model. Then, the intermediate frame is obtained by using pixel warping and fusion strategies.

[0090] Step 4) Compare the synthesized intermediate frame with the second frame of the triplet, and use gradient descent to update the parameters in the motion refinement model; iterate repeatedly until the optimal parameter combination is trained to obtain the trained motion refinement model.

[0091] As an improvement to the above technical solution, in step 4), the parameters in the motion refinement model are updated based on the output of the motion refinement model and the loss of the true label. The loss function is the Charbonnier loss, and its expression is:

[0092]

[0093] in, I and Hwc represent the model output and the true label, respectively, while ε represents the set of pixels. 2 It is a very small positive number.

[0094] The present invention also proposes a motion refinement system for video frame interpolation, the system comprising:

[0095] The preprocessing module processes the two frames to be interpolated to obtain the raw optical flow estimate between the two frames, which serves as the optical flow feature; and

[0096] The intermediate frame acquisition module is used to input two frames of images and the original optical flow estimate into a pre-established and trained motion thinning model to achieve motion thinning, obtaining an improved optical flow estimate and pixel reliability score; and based on the improved optical flow estimate and pixel reliability score, an intermediate frame is synthesized through pixel warping and fusion strategies; wherein...

[0097] The motion refinement model includes: a downsampling module, a context feature extraction module, a joint stream coding module, a motion-guided feature fusion module, and a decoder; wherein...

[0098] The downsampling module uses two image pyramid structures to downsample the features of two frames of images respectively.

[0099] The context feature extraction module is used to enhance the image features of each layer of the downsampling module;

[0100] The joint stream coding module is used to interactively fuse image features and optical flow features;

[0101] The motion-guided feature fusion module is used to fuse pixel information based on prior knowledge of optical flow estimation and to process the features fused by the joint stream coding module to achieve fine fusion of guided optical flow features and image features.

[0102] The decoder is used to decode optical flow features and image features obtained from the joint stream coding module and the feature fusion module, and to generate an improved optical flow estimate and its corresponding pixel reliability score.

[0103] The advantages of this invention compared to the prior art are:

[0104] 1. Based on existing improved optical flow estimation methods, this invention proposes two modules to further enhance optical flow estimation: First, a contextual feature extraction module is embedded in each layer of the image feature pyramid, which can effectively extract global and local information of the image; second, the main network needs to fuse bilateral image and optical flow features. The fusion strategy of multi-source features needs to consider not only the segmentation of moving objects and boundary static objects, but also the internal information of the fused objects. To achieve this, this invention proposes a motion-guided feature fusion module, which fuses pixel information based on the prior knowledge of existing optical flow estimation. Since existing optical flow estimation provides an approximate estimate of the motion field of moving objects, the fusion strategy based on optical flow estimation can automatically adjust the fusion of pixels at different positions of the moving object, thereby making the separation between the moving object and the boundary object clearer, while also improving the fidelity of the generated moving object, making it closer to the original image.

[0105] 2. The Comprehensive Contextual Feature Extraction (CCFE) module is integrated into each layer of the pyramid structure. It aims to encourage the model to extract clean and sufficiently rich contextual information from the input image. The Motion-Guided Feature Fusion (MGFF) module can guide the fusion of image features based on optical flow features, making the feature fusion of moving objects more accurate, thereby providing local compensation for optical flow estimation. Through the above two modules, this invention can obtain enhanced motion optical flow estimation. This optical flow estimation describes the inter-frame motion field information in detail, which can effectively handle the occlusion and pixel blur problems in large motion scenes and achieve high-quality video frame interpolation. Attached Figure Description

[0106] Figure 1 A schematic diagram of a motion thinning network structure with local compensation for video frame interpolation provided by the present invention;

[0107] Figure 2(a) shows the structure of the multi-scale feature aggregation mechanism, and Figure 2(b) shows the structure of the channel attention mechanism block.

[0108] Figure 3 A schematic diagram of the joint stream coding module structure provided by this invention;

[0109] Figure 4 This is a schematic diagram of the feature fusion module structure for motion guidance provided by the present invention. Detailed Implementation

[0110] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0111] Example 1

[0112] The purpose of this invention is to overcome the technical defects of existing video methods and propose a motion thinning method with local compensation for video frame interpolation, specifically including:

[0113] Step 1) Obtain data: Randomly read a set of triplet data. Each triplet contains three consecutive frames of images. The first and third frames are the input to the motion refinement model, and the second frame is the real label for supervised training.

[0114] Step 2) Obtain the original optical flow estimate. Input the first and third frames I0 and I1 from the triplet image into the existing optical flow estimation model PWC-NET, and estimate the original optical flow F in two directions between the two input frames. 0->1 and F 1->0 ;

[0115] Step 3) Take the first and third frames I0 and I1 from the triplet image, and the original optical flow estimate F 0->1 and F 1->0 The improved optical flow estimate and pixel reliability score are obtained by inputting the data into the established motion refinement model. Then, the intermediate frame is obtained through pixel warping and fusion strategies.

[0116] Step 4) Compare the synthesized intermediate frame with the second frame of the triplet, and use gradient descent to update the parameters in the motion refinement model; iterate repeatedly until the optimal parameter combination is trained to obtain the trained motion refinement model.

[0117] Furthermore, such as Figure 1 As shown, step 3) specifically includes:

[0118] Step 3-1) The input image is downsampled using a downsampling module. Each downsampling module is implemented using two two-dimensional convolutions with different parameters and a PReLU activation function; the image features of the l-th downsampling module are... sl The expression is:

[0119] I sl =PReLu(Conv2d(PReLu(Cond2d(I s(l-1) ))))

[0120] Where s∈{0,1} represents the frame source, l∈{1,…,L} represents the first layer, L represents the total number of layers, and I s(l-1) The image features are represented by the (l-1)th downsampling layer. Cond2d() represents a two-dimensional convolution, and PReLu() represents the activation function. In this embodiment, for ease of understanding, L = 4 is used; in practice, the value of L can be flexibly set according to requirements.

[0121] Step 3-2) Employs a context feature extraction module to enhance the features obtained by the downsampling module. Each context feature extraction module includes a multi-scale feature aggregation mechanism block and a channel attention mechanism block. The processing procedure specifically includes:

[0122] Step A1. Convert the feature map I obtained by the downsampling module sl The input multi-scale feature aggregation mechanism block is shown in Figure 2(a). The multi-scale feature aggregation mechanism consists of k dilated convolution branches with different dilation rates. The k dilated convolution branches respectively process image features I. sl The process involves generating new feature maps, and the steps are as follows:

[0123] O k =dilated conv k (I sl )

[0124] Among them, dilated conv k () and O k Let represent the dilated convolution operation and output of the k-th branch, respectively;

[0125] Step A2. Concatenate the features generated by all branches along the channel dimension and pass them through the δ activation function to obtain a multi-scale feature map u. sl The expression is:

[0126] u sl =δ(Concat(O) k ))

[0127] Where δ() represents the ReLU activation function, and Concat represents concatenation;

[0128] Step A3. Put u sl The input channel attention mechanism block, as shown in Figure 2(b), includes a global average pooling layer, two convolutional layers, and a multiplication operation layer.

[0129] The global average pooling layer performs a global average pooling operation on the feature map u. sl Compressed along the spatial dimension into a feature vector z csl The process expression is:

[0130]

[0131] Where G() represents the global average pooling operation, H×W represents the size of the input feature map, W represents the width, and H represents the height; u sl (x, y) represents the feature map u sl The eigenvalue at a height of x and a width of y;

[0132] Step A4. Transfer feature z csl The vector is transformed into learnable parameters, and the specific formula is as follows:

[0133] v csl =T(z) csl ,w)=σ(Conv2(δ(Conv1(z csl )))

[0134] Where T(·) represents the vector transformation function, δ represents the ReLU activation function, and σ represents the sigmoid activation function; Conv1 and Conv2 are respectively the transformation functions of the feature vector z. csl One-dimensional and two-dimensional convolutions are used for mapping;

[0135] Step A4. The product operation layer weights the learned feature channel importance coefficients w with their corresponding feature channels to obtain the recalibrated features x. sl The formula is:

[0136] x sl =v csl ·u csl ,where c∈(0,1,…,C)

[0137] Where c and C represent the current number of channels and the total number of channels, respectively, v csl u represents the channel weight calculated at channel c in the l-th layer of image s. csl This represents the feature value of image s at channel c in the l-th layer.

[0138] Step 3-3) Use a joint stream coding module to distort the optical flow features and image features, such as... Figure 3 As shown, the processing procedure specifically includes:

[0139] Step B1. From layer 1 to layer L, apply the joint stream coding module to generate the optical flow feature pyramid of the bidirectional flow field;

[0140] In the l-th layer of the joint stream coding module, the optical flow features and image features from the previous layer are distorted together. Specifically, the optical flow estimation from the previous layer is used, along with the corresponding pyramid features I. 0(l-1) To I 1(l-1) Distortion, obtaining the distortion feature wim 1l ; and the corresponding pyramid feature I 1(l-1) To I 0(l-1) Distortion, obtaining the distortion feature wim 0l The expressions are as follows:

[0141] wim 1l =warp(I 1(l-1) f 0(l-1) )

[0142] wim 0l =warp(I 0(l-1) f 1(l-1) )

[0143] Among them, warp() means warp;

[0144] Step B2. Concatenate and combine the original optical flow features, image features, and distortion features, and then downsample them using two convolutional layers:

[0145] e 0l =Down(Concat(f 0(l-1) I 0(l-1) wim 1l ))

[0146] e 1l =Down(Concat(f 1(l-1) I 1(l-1) wim 0l ))

[0147] Concat() represents concatenation, and Down() represents downsampling;

[0148] Step B3. Encode the optical flow features of layer l (l-1) to obtain the optical flow features f of layer l on the two branches. 0l and f 1l The expressions are as follows:

[0149] f 0l =Down(f 0(l-1) )

[0150] f 1l =Down(f 1(l-1) ).

[0151] Steps 3-4) use a motion-guided feature fusion module to fuse the hybrid features e based on the Lth layer output of the joint stream coding module. L Optical flow characteristics fL, image downsampling characteristics I L ,like Figure 4 As shown, the specific processing steps include:

[0152] Step C1. The feature fusion module uses the hybrid feature e output from the Lth layer of the joint stream coding module. L Optical flow characteristics f L Image downsampling features I L After fusion, three multi-scale weighted features w1, w2, and w3 are obtained, expressed as follows:

[0153]

[0154]

[0155]

[0156] in, and Indicates a fully connected layer operation, Down p () represents the downsampling function implemented by a 2D convolution with a stride of 2, where p1 and p2 represent the downsampling ratios;

[0157] Step C2. Merge the weighted features to obtain p1 times upsampling weight w4 and p2 times upsampling weight w o :

[0158] w4 = w2 + Up d (w3)

[0159]

[0160] Among them, Up d () denotes the upsampling function implemented by bilinear interpolation, where d is the sampling factor; For a fully connected layer, σ() is the Sigmoid activation function;

[0161] Step C3. Obtain the fused feature e using the following formula. l+1 :

[0162] e L =e L ·w o +I L ·(1-w o ).

[0163] Steps 3-5) use an encoder to obtain improved optical flow estimation and pixel reliability scores. The specific processing steps include:

[0164] Step D1. The decoder obtains feature maps of the optical flow feature pyramid and the image feature pyramid based on the joint stream coding module and the feature fusion module, and generates N motion sub-vectors step by step using deconvolution. And its corresponding pixel reliability scores {S0, S1}, where m0, m1, m2 and m3 represent hidden layer features, calculated as follows:

[0165] m0 = PReLu(DeConv(Concat[e L c L f L ]))

[0166] m1 = PReLu(DeConv(Concat[e L-1 ,m0]))

[0167] m2=PReLu(DeConv(Concat[e L-2 ,m1]))

[0168] m3=PReLu(DeConv(Concat[e L-3 ,m2]))

[0169]

[0170] Where PReLu() represents the activation function and DeConv() represents deconvolution;

[0171] Steps 3-6) employ pixel warping and fusion strategies to obtain improved multi-motion vectors based on the decoder. The intermediate frames are synthesized from {S0, S1} using the following steps:

[0172] Step D2. First, scale the multiple motion vectors to the given target time step t:

[0173]

[0174]

[0175] Where i0 and i1 represent the i-th source pixel of I0 and I1, respectively; This represents the optical flow estimate of the nth motion vector at pixel i0 from source frame 0 to source frame 1. This represents the optical flow estimate of the nth motion vector at pixel i1 from source frame 1 to source frame 0. Indicates will Optical flow estimate scaled to target time t Indicates will Optical flow estimate scaled to target time t;

[0176] Step D3. Then, a source pixel i s The nth motion vector will be warped forward by the forward warping operation φ to the source pixel i s Warped The expression is:

[0177]

[0178] Step D4. The nth motion sub-vector can warp each pixel in the source frame s forward to the target time t, thus obtaining... The set of target pixels after warping of N sub-motion vectors:

[0179]

[0180] Step D5. Measure the importance of each pixel from three aspects: temporal correlation, brightness consistency, and reliability score:

[0181] Time correlation is represented by r i Let r represent the value of i when i is a pixel from I0. i = 1-t, and when i is a pixel from I1, r i =t;

[0182] Brightness uniformity is represented by b i It is indicated that it is obtained in the following way:

[0183]

[0184] The reliability scores {S0, S1} are obtained through step D1;

[0185] Step D6. Using pixel warping and fusion strategies, the original pixel i is fused to obtain the intermediate frame pixel j:

[0186]

[0187] Among them, s i It is the pixel reliability score at position i; c i α is the original color of pixel i; α is a learnable parameter used to adjust the importance of the weights.

[0188] Furthermore, step 4) specifically includes:

[0189] The parameters in the motion refinement model are updated based on the output of the motion refinement model and the loss of the true labels. The loss function is the Charbonnier loss, expressed as:

[0190]

[0191] in, I and Hwc represent the model output and the true label, respectively, while ε represents the set of pixels. 2 It is a very small positive number.

[0192] Example 2

[0193] Embodiment 2 of the present invention provides a motion refinement system for video frame interpolation, the system comprising:

[0194] The preprocessing module processes the two frames to be interpolated to obtain the raw optical flow estimate between the two frames, which serves as the optical flow feature; and

[0195] The intermediate frame acquisition module is used to input two frames of images and the original optical flow estimate into a pre-established and trained motion thinning model to achieve motion thinning, obtaining an improved optical flow estimate and pixel reliability score. Based on the improved optical flow estimate and pixel reliability score, intermediate frames are synthesized through pixel warping and fusion strategies.

[0196] The motion refinement model includes: a downsampling module, a context feature extraction module, a joint stream coding module, a motion-guided feature fusion module, and a decoder; wherein...

[0197] The downsampling module uses two image pyramid structures to downsample the features of two frames of images respectively.

[0198] The context feature extraction module is used to enhance the image features of each layer of the downsampling module;

[0199] The joint stream coding module is used to interactively fuse image features and optical flow features;

[0200] The motion-guided feature fusion module is used to fuse pixel information based on prior knowledge of optical flow estimation and to process the features fused by the joint stream coding module to achieve fine fusion of guided optical flow features and image features.

[0201] The decoder is used to decode optical flow features and image features obtained from the joint stream coding module and the feature fusion module, and to generate an improved optical flow estimate and its corresponding pixel reliability score.

[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the embodiments, those skilled in the art should understand that modifications or equivalent substitutions to the technical solutions of the present invention do not depart from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A motion refinement method for video interpolation, the method comprising: processing two frames of images to be interpolated to obtain an original optical flow estimation between the two frames of images as an optical flow feature; inputting the two frames of images and the original optical flow estimation into a pre-established and trained motion refinement model to achieve motion refinement, to obtain an improved optical flow estimation and a pixel reliability score; based on the improved optical flow estimation and the pixel reliability score, synthesizing an intermediate frame through a pixel warping and fusion strategy; wherein, the motion refinement model comprises: a downsampling module, a context feature extraction module, a joint flow encoding module, a motion-guided feature fusion module, and a decoder; wherein, the downsampling module uses two image pyramid structures to downsample the features of the two frames of images, respectively; the context feature extraction module is configured to enhance the image features of each layer of the downsampling module; the joint flow encoding module is configured to interactively fuse the image features and the optical flow features; the motion-guided feature fusion module is configured to fuse pixel information based on the prior knowledge of the optical flow estimation, and process the features interactively fused by the joint flow encoding module to achieve fine fusion of the guided optical flow features and the image features; the decoder is configured to decode the optical flow features and the image features obtained by the joint flow encoding module, the context feature extraction module, and the feature fusion module, and generate an improved optical flow estimation and a corresponding pixel reliability score; the method of synthesizing an intermediate frame based on the improved optical flow estimation and the pixel reliability score through a pixel warping and fusion strategy comprises the following steps: Step D2. First, the multiple motion vectors are scaled to the given target time step t: ; ; wherein, and respectively denote and the n-th motion vector of source pixel ; denotes the optical flow estimate of the n-th motion vector of source frame 0 to source frame 1 at pixel , denotes the optical flow estimate of the n-th motion vector of source frame 1 to source frame 0 at pixel , denotes the optical flow estimate of scaling to the target time t, denotes the optical flow estimate of scaling to the target time t; Step D3. Then, a source pixel The The motion vectors are warped forward. The source pixel Warped The expression is: ; Step D4. The nth motion sub-vector warps each pixel in the source frame s forward in time to the target time t, resulting in The set of warped target pixels is: ; Step D5. Measure the importance of each pixel from three aspects of temporal correlation, brightness consistency, and reliability score: Time correlation is expressed by when i is a pixel point from , and when i is a pixel point from , ; The brightness uniformity is expressed by the following equation: ; Reliability score obtained by step D1 ; Step D6. The original pixels are fused to get the intermediate frame pixels by a pixel warping and fusing strategy i Step D6. The original pixels are fused to get the intermediate frame pixels by a pixel warping and fusing strategy j : ; wherein, is the pixel reliability score for the i th position; is the original color of the pixel i ; and is a learnable parameter that adjusts the importance of the weight.

2. The method for motion refinement for video interpolation according to claim 1, wherein, The downsampling modules are both implemented by two-dimensional convolution with two different parameters and PReLu activation function; the first downsampling module is implemented by two-dimensional convolution with a 3*3 kernel and PReLu activation function, and the second downsampling module is implemented by two-dimensional convolution with a 1*1 kernel and PReLu activation function. l The image features of the two downsampling modules The expression is: )); wherein denotes a frame source, denotes the l-th layer, denotes the total number of layers, is the l-th l -1 down-sampling layer of image features, denotes a two-dimensional convolution, denotes an activation function.

3. The method for motion refinement for video interpolation according to claim 2, wherein, the context feature extraction module comprises a multi-scale feature aggregation mechanism block and a channel attention mechanism block, and the processing process specifically comprises: Step A1. The feature map obtained by the down-sampling module The input multi-scale feature aggregation mechanism block includes The first to the third convolutional branches with different hole rates are constructed, The first to the third convolutional branches respectively process the image features to generate new feature maps, and the processing process is represented as: ; wherein, ( ) and respectively represent the hollow convolution operation and the output of the i-th branch. th branch. Step A2. Concatenate all branch-generated features in the channel dimension and pass them through an activation function to obtain multi-scale feature maps , expressed as: ; wherein, ( ) denotes a ReLu activation function, denotes concatenation; Step A3. The compound of formula (I) is prepared by the following reaction scheme: an input channel attention mechanism block comprising one global average pooling layer, two convolution layers and one product operation layer; The global average pooling layer performs a global average pooling operation to compress the feature map into one feature vector along the spatial dimension The process expression is: ; wherein, denotes a global average pooling operation, denotes a size of an input feature map, denotes a width, H denotes a height; denotes a feature map at a height of , a width of a feature value; Step A4. The features The vector is converted into learnable parameters with the formula: ; wherein, represents a vector transformation function, represents a ReLu activation function, represents a sigmoid activation function; and are a one-dimensional and two-dimensional convolution, respectively, mapping a feature vector ​ Step A4. The product operation layer will weight the learned feature channel importance coefficients with their corresponding feature channels to get re-targeted features , which is given by: ; wherein, c and C C and Ctotal represent the current and total number of channels, respectively, represents an image In the first layer channel is the channel weight calculated at the represents an image In the first layer channel is the feature value at the 4. The method for motion refinement for video interpolation according to claim 3, wherein, the processing process of the joint flow encoding module specifically comprises: Step B1. From the 1st layer to the 2nd layer, apply the joint flow encoding module to generate the optical flow feature pyramid of the bidirectional flow field. Step B1. From the 1st layer to the 2nd layer, apply the joint flow encoding module to generate the optical flow feature pyramid of the bidirectional flow field. In the joint stream coding module Within the layer, the optical flow features and image features of the previous layer are distorted, specifically: the optical flow estimation from the previous layer is compared with the features of the corresponding pyramid. Towards Distortion, to obtain distorted features ; and the characteristics of the corresponding pyramid Towards Distortion, to obtain distorted features The expressions are as follows: ; ; wherein represents a twist; Step B2. Splice and combine the original feature optical flow features, image features, and distortion features, and use two layers of convolution to downsample them: ; ; wherein ( ) denotes concatenation, denotes down-sampling; Step B3. Downsample the optical flow features of the first layer to get two branches of the first layer optical flow features and , whose expressions are respectively: ; 。 5. The method for motion refinement for video interpolation according to claim 4, wherein, the processing process of the motion-guided feature fusion module specifically comprises: Step C1. The feature fusion module fuses the features based on the joint stream encoding module layer output mixed features , optical flow features , image down-sampling features to obtain three multi-scale weight features , and , and the expression is: ; ; ; wherein, ( ), ( ) and ( ) denotes a fully connected layer operation, ( ) denotes a down-sampling function implemented by a two-dimensional convolution with a step of 2, and denotes the ratio of down-sampling; Step C2. The weight features are combined to obtain up-sampling weights and up-sampling weights : ; ; wherein, ( ) represents a bilinear interpolation implemented up-sampling function, wherein is a sampling factor; ( ) is a fully connected layer, ( ) is a Sigmoid activation function; Step C3. The resulting features are fused by : 。 6. The method for motion refinement for video interpolation according to claim 5, wherein, the processing process of the decoder specifically comprises: Step D1. The decoder obtains the feature maps of the optical flow feature pyramid and the image feature pyramid based on the joint stream encoding module and the feature fusion module, and generates N motion sub-vectors step by step by using deconvolution That is, the improved optical flow estimation and the corresponding pixel reliability score , 、 、 and represent the hidden layer features, and the calculation formula is: PReLu(DeConv(Concat[ ]))); PReLu(DeConv(Concat[ ]))); PReLu(DeConv(Concat[ ]))); PReLu(DeConv(Concat[ ]))); = Conv( ) ; wherein, PReLu ( ) represents an activation function, and DeConv( ) represents deconvolution.

7. The method for motion refinement for video inter- framing of claim 1, wherein, The method further comprises training the motion refinement model to obtain a trained motion refinement model, and the training process specifically comprises: Step 1) Obtain data, randomly read a set of three-tuple data, each three-tuple contains three consecutive frames of images, the first and third frames of images are the inputs of the motion refinement model, and the second frame of image is the real label for supervised training; Step 2) Obtain the raw optical flow estimate by combining the first and third frames of the triplet image. and The input is fed into the existing optical flow estimation model PWC-NET to estimate the original optical flow in two directions between two input frames. and ; Step 3) Extract the first and third frames from the triplet image. and Original optical flow estimation and The improved optical flow estimate and pixel reliability score are obtained by inputting the data into the established motion refinement model. Then, the intermediate frame is obtained by using pixel warping and fusion strategies. Step 4) Compare the synthesized intermediate frame with the second frame of the three-tuple, and update the parameters in the motion refinement model using the gradient descent method; repeatedly iterate until the optimal parameter combination is trained, to obtain the trained motion refinement model.

8. The method for motion refinement for video interpolation according to claim 7, wherein, In the step 4), the parameters in the motion refinement model are updated according to the loss of the output of the motion refinement model and the real label, and the loss function is Charbonnier loss, and the expression is: ; wherein, respectively denote the model output and the true label, denotes a set of pixels, is a very small positive number.

9. A system for motion refinement for video interpolation based on the method of claim 1, characterized by, The system comprises: The pre-processing module is configured to process two frames of images to be interpolated to obtain an original optical flow estimation between the two frames of images as an optical flow feature; and The intermediate frame acquisition module is configured to input the two frames of images and the original optical flow estimation into a motion refinement model that is pre-established and trained, to implement motion refinement to obtain an improved optical flow estimation and a pixel reliability score; and based on the improved optical flow estimation and the pixel reliability score, to synthesize an intermediate frame through a pixel warping and fusion strategy; wherein The motion refinement model comprises a down-sampling module, a context feature extraction module, a joint flow encoding module, a motion-guided feature fusion module, and a decoder; wherein The down-sampling module is configured to down-sample two frame image features using two image pyramid structures, respectively. The context feature extraction module is configured to enhance each layer of image features of the down-sampling module. The joint flow encoding module is configured to interactively fuse the image features and the optical flow features. The motion-guided feature fusion module is configured to fuse pixel information based on prior knowledge of the optical flow estimation, to process the features interactively fused by the joint flow encoding module to implement fine fusion of the guided optical flow features and the image features. The decoder is configured to decode the optical flow features and the image features obtained by the joint flow encoding module and the feature fusion module, and to generate an improved optical flow estimation and a corresponding pixel reliability score.

Citation Information

Patent Citations

  • Method and device for generating video intermediate frame

    CN115065796A

  • End-to-end global and local motion estimation method based on deep learning

    CN116091555A