A Bidirectional Encoding Video Frame Interpolation Method, System and Device Based on Deep Learning

Through the combination of the bidirectional encoder structure and the channel attention cascade module, the problem of high complexity of the existing CNN model is solved, and high-quality video interpolation effect is achieved.

CN114125455BActive Publication Date: 2025-07-11LUOYANG NO 5 GONGGUAN CULTURE MEDIA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111394443.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-23
Publication Date
2025-07-11
Estimated Expiration
2041-11-23

AI Technical Summary

Technical Problem

The existing CNN-based video interpolation method model design is complex, resulting in large model scale, high computational complexity, and low video interpolation quality.

Method used

Using a bidirectional encoder structure, combining the standard convolution of low-dimensional feature layers and the depth of high-dimensional feature layers can be separated convolution, and adding a channel attention cascade module to reduce the computational complexity through the fusion of forward and backward coding features using local lightweighting.

Benefits of technology

It improves the quality of video interpolation, reduces the computational complexity, and enhances the integrity and richness of feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114125455B_ABST
    Figure CN114125455B_ABST
Patent Text Reader

Abstract

The present invention discloses a two-way encoding video frame interpolation method, system and device based on deep learning. The method obtains two adjacent frames in a target video, uses a first encoder to perform forward encoding on the two adjacent frames to obtain first encoded features, uses a second encoder to perform backward encoding on the two adjacent frames to obtain second encoded features, fuses the first encoded features and the second encoded features to obtain fused encoded features; uses a decoder to decode the fused encoded features to obtain decoded features; performs convolution on the decoded features to obtain convolution parameters; and calculates and obtains a target synthesized frame based on the convolution parameters and the two adjacent frames. This application uses two encoders to extract features from the input frames and fuse the extracted features, making the extracted features more complete and rich, and enhancing the video frame interpolation quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and in particular, to a bidirectional encoding video frame interpolation method, system and device based on deep learning. Background Art

[0002] With the rapid application of CNN in video frame interpolation, the latest video frame interpolation methods generally use CNN for frame interpolation, and there is an increasing trend to design complex and heavy models based on CNN for video frame interpolation. The design of such CNN-based models is becoming increasingly complex, resulting in large model sizes and high computational complexities. For example, Lee, Hyeongmin, et al. "Adacof: Adaptive collaboration of flows for video frame interpolation." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020: 5316-5325. proposed a network structure including a U-Net, a sub-network, and an AdaCoF module. Only one encoder and one decoder were used in the U-Net, and the features extracted using one encoder were not complete and rich enough. Therefore, the quality of video frame interpolation is not high. Summary of the Invention

[0003] To solve the problems existing in the prior art, the present invention provides a bidirectional encoding video frame interpolation method, system and device based on deep learning, which can ensure the quality of video frame interpolation.

[0004] To achieve the above object, the technical solution of the present invention is implemented as follows:

[0005] In a first aspect, the present invention provides a bidirectional encoding video frame interpolation method based on deep learning, including the following steps:

[0006] Obtain two adjacent frames in a target video, perform forward encoding on the two adjacent frames using a first encoder to obtain a first encoded feature, perform backward encoding on the two adjacent frames using a second encoder to obtain a second encoded feature, and fuse the first encoded feature and the second encoded feature to obtain a fused encoded feature;

[0007] Use a decoder to decode the fused encoded feature to obtain a decoded feature;

[0008] Perform convolution on the decoded feature to obtain convolution parameters;

[0009] Based on the convolution parameters and the two adjacent frames, calculate to obtain a target synthesized frame.

[0010] Compared with the prior art, the present invention has the following beneficial effects:

[0011] In the present application, the first encoder performs forward encoding on the two adjacent frames to obtain a first encoded feature, the second encoder performs backward encoding on the two adjacent frames to obtain a second encoded feature, and the first encoded feature and the second encoded feature are fused, making the extracted features more complete and rich, and enhancing the video frame interpolation quality.

[0012] Further, the step of using the first encoder to perform forward encoding on the two adjacent frames to obtain a first encoded feature and using the second encoder to perform backward encoding on the two adjacent frames to obtain a second encoded feature includes:

[0013] Using the standard convolution of the low-dimensional feature layer and the depthwise separable convolution of the high-dimensional feature layer in the first encoder to perform the forward encoding on the motion information of two adjacent frames of the multiple input frames from the previous frame to the next frame to obtain the first encoded feature;

[0014] Using the standard convolution of the low-dimensional feature layer and the depthwise separable convolution of the high-dimensional feature layer in the second encoder to perform the backward encoding on the motion information of two adjacent frames of the multiple input frames from the next frame to the previous frame to obtain the second encoded feature.

[0015] Further, the step of using the decoder to decode the fused encoded feature to obtain a decoded feature includes:

[0016] Using the standard convolution of the low-dimensional feature layer and the depthwise separable convolution of the high-dimensional feature layer in the decoder to decode the fused encoded feature to obtain the decoded feature.

[0017] Further, it further includes:

[0018] A channel attention cascade module is added between the first encoder and the decoder and between the second encoder and the decoder.

[0019] Further, the step of performing convolution on the decoded feature to obtain convolution parameters includes:

[0020] Using multiple groups of standard convolutions to perform convolution operations on the decoded feature to obtain the convolution parameters.

[0021] Further, the step of calculating and obtaining a target synthesized frame based on the convolution parameters and the two adjacent frames includes:

[0022] Based on the convolution parameters and the two adjacent frames, the target synthesized frame is calculated through the following calculation formula, and the calculation formula is:

[0023]

[0024] I t (i, j) = V ⊙ I n (i, j) + (1 - V) ⊙ I n+1 (i, j) (2)

[0025] In Equation (1), I(i, j) represents the value of the pixel point obtained after the input frame is calculated by AdaCoF, K represents the size of the kernel, i represents the x-axis coordinate of the pixel point, j represents the y-axis coordinate of the pixel point, α k,l represents the vertical offset of the pixel point of the input frame, β k,l represents the horizontal offset of the pixel point of the input frame, W k,l represents the weight of the kernel, I(i + dk + α k,l , j ± dl + β k,l ) represents the value of the pixel point of the input frame, dk represents the vertical dilation of the pixel point of the input frame, and dl represents the horizontal dilation of the pixel point of the input frame;

[0026] In Equation (2), V represents the weight obtained by a group of convolutions, the value range of V is from 0 to 1, ⊙ represents element-wise multiplication, and I n (i, j) represents the value of the pixel point obtained after the previous frame of two adjacent frames is calculated by Equation (1), and I n+1 (i, j) represents the value of the pixel point obtained after the subsequent frame of two adjacent frames is calculated by Equation (1), and I t (i, j) represents the value of the pixel point of the target synthesis frame.

[0027] In a second aspect, the present invention provides a two-way encoding video frame interpolation system based on deep learning, including:

[0028] A fusion encoding feature acquisition unit, configured to acquire two adjacent frames in a target video, perform forward encoding on the two adjacent frames using a first encoder to obtain a first encoding feature, perform backward encoding on the two adjacent frames using a second encoder to obtain a second encoding feature, and fuse the first encoding feature and the second encoding feature to obtain a fusion encoding feature;

[0029] A decoded feature acquisition unit, configured to decode the fusion encoding feature using a decoder to obtain a decoded feature;

[0030] A parameter acquisition unit, configured to perform convolution on the decoded feature to obtain convolution parameters;

[0031] A target synthesis frame obtaining unit, configured to calculate and obtain a target synthesis frame based on the convolutional parameters and the two adjacent frames.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] In the present application, the fusion coding feature obtaining unit uses a first encoder to perform forward coding on the two adjacent frames to obtain first coding features, uses a second encoder to perform backward coding on the two adjacent frames to obtain second coding features, and fuses the first coding features and the second coding features, making the extracted features more complete and rich, and enhancing the video frame interpolation quality.

[0034] In a third aspect, a deep learning-based bidirectional coding video frame interpolation device includes at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and when the instructions are executed by the at least one control processor, the at least one control processor is enabled to execute a deep learning-based bidirectional coding video frame interpolation method as described above.

[0035] In a fourth aspect, the present invention provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute a deep learning-based bidirectional coding video frame interpolation method as described above. Description of the Drawings

[0036] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:

[0037] Figure 1 is a flowchart of a deep learning-based bidirectional coding video frame interpolation method provided by an embodiment of the present invention;

[0038] Figure 2 is a structure diagram of a prior network related to the present invention;

[0039] Figure 3 is a structure diagram of a deep learning-based bidirectional coding video frame interpolation method provided by an embodiment of the present invention;

[0040] Figure 4 is a structure diagram of a standard convolution provided by an embodiment of the present invention;

[0041] Figure 5 is a structure diagram of a depthwise separable convolution provided by an embodiment of the present invention;

[0042] Figure 6Structural diagram of the channel attention cascading module provided by an embodiment of the present invention;

[0043] Figure 7 Structural diagram of a two-way encoded video frame interpolation system based on deep learning provided by an embodiment of the present invention. Detailed implementation manners

[0044] Next, the technical solutions of the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other. In addition, the accompanying drawings are used to supplement the description of the text part of the specification, enabling people to intuitively and vividly understand each technical feature and the overall technical solution of the present disclosure, but they should not be construed as limiting the scope of protection of the present disclosure.

[0045] In the description of the present invention, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0046] In the description of the present invention, unless otherwise clearly defined, terms such as "set", "installed", and "connected" should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meanings of the above terms in the present invention in combination with the specific content of the technical solution.

[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0048] With the rapid application of Convolutional Neural Network (CNN) in video frame interpolation, the latest video frame interpolation methods generally use CNN for frame interpolation, and there is an increasing trend to design complex and heavy models based on CNN for video frame interpolation. The design of such CNN-based models is becoming increasingly complex, resulting in large model sizes and high computational complexities. For example, Lee, Hyeongmin, et al. "Adacof: Adaptive collaboration of flows for video frame interpolation." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020: 5316-5325. proposed a network structure, such as Figure 2 described, including a U-Net, a sub-network, and an AdaCoF module. Only one encoder and one decoder were used in the U-Net, and the features extracted using only one encoder were not complete and rich enough. Therefore, the quality of video frame interpolation was not high.

[0049] To solve the above problems, compared with Figure 2 , this application uses two encoders on the above network structure to generate a new network structure. This network structure uses the two encoders to extract features from the input frames and fuse the extracted features, making the extracted features more complete and rich, and enhancing the quality of video frame interpolation.

[0050] Referring to Figure 1 , an embodiment of the present invention provides a bidirectional encoding video frame interpolation method based on deep learning, including the following steps:

[0051] Step S100: Obtain two adjacent frames in the target video, perform forward encoding on the two adjacent frames using a first encoder to obtain a first encoded feature, perform backward encoding on the two adjacent frames using a second encoder to obtain a second encoded feature, and fuse the first encoded feature and the second encoded feature to obtain a fused encoded feature.

[0052] Furthermore, due to the increasingly complex design of the CNN model, resulting in a large model size and high computational complexity, and the above first encoder and second encoder are both designed based on the CNN model, the computational complexities of the first encoder and the second encoder are high. To reduce the computational complexity and ensure the quality of video frame interpolation in this embodiment, referring to Figure 3, local lightweighting is used for the first encoder and the second encoder above. The local lightweighting refers to replacing the standard convolution of the high-dimensional feature layer with depthwise separable convolution, as Figure 4 and Figure 5 described, Figure 4 is the structure of the standard convolution, Figure 5 is the structure of the depthwise separable convolution.

[0053] Specifically, the standard convolution of the low-dimensional feature layer and the depthwise separable convolution of the high-dimensional feature layer in the first encoder are used to forward-encode the motion information of two adjacent frames of multiple input frames from the previous frame to the next frame, obtaining the first encoded feature.

[0054] The standard convolution of the low-dimensional feature layer and the depthwise separable convolution of the high-dimensional feature layer in the second encoder are used to backward-encode the motion information of two adjacent frames of multiple input frames from the next frame to the previous frame, obtaining the second encoded feature.

[0055] Then, the first encoded feature and the second encoded feature are concatenated to obtain the fused encoded feature.

[0056] Among them, Figure 3 (SeparableConv+ReLU)×3 in represents the convolution in the high-dimensional feature layer that uses depthwise separable convolution in the two encoders, and (Conv+ReLU)×3 represents the convolution in the low-dimensional feature layer that uses standard convolution in the two encoders. Among them, SeparateConv represents depthwise separable convolution, Conv represents standard convolution, and ReLU represents the activation function.

[0057] For better illustration, the following formula is used for illustration:

[0058]

[0059] Among them, L encoder represents the encoding process, F forward represents the first encoded feature obtained forward, F backward represents the second encoded feature obtained backward, represents two adjacent frames for forward encoding, represents two adjacent frames for backward encoding, [F forward ,F backward represents the fused encoded feature obtained after concatenating the first encoded feature and the second encoded feature.

[0060] In this embodiment, two encoders are used to extract features from the input frame and fuse the extracted features, making the extracted features more complete and rich, and improving the quality of video frame interpolation. Local lightweighting is also used to reduce the number of parameters, lower the computational complexity, and ensure the quality of video frame interpolation.

[0061] Step S200: Use a decoder to decode the fused encoded features to obtain decoded features.

[0062] Specifically, since the decoder is also designed based on the CNN model, the computational complexity of the decoder is also high. Referring to the method in step S100, in this embodiment, the standard convolution of the high-dimensional feature layer in the decoder is replaced with depthwise separable convolution. The standard convolution of the low-dimensional feature layer in the decoder and the depthwise separable convolution of the high-dimensional feature layer are used to decode the fused encoded features obtained by concatenating the first encoded feature and the second encoded feature to obtain decoded features. Among them, Figure 3 (Upsample+SeparableConv+ReLU) represents the convolution in the high-dimensional feature layer using depthwise separable convolution in the decoder, and (Upsample+Conv+ReLU) represents the convolution in the low-dimensional feature layer using standard convolution in the decoder, where SeparateConv represents depthwise separable convolution, Conv represents standard convolution, ReLU represents the activation function, and Upsample represents upsampling.

[0063] For better illustration, the following formula is used for explanation:

[0064] O = L decoder ([F forward ,F backward )

[0065] Among them, L decoder represents the decoding process, O represents the obtained decoded features, and [F forward ,F backward represents the fused encoded features obtained by concatenating the first encoded feature and the second encoded feature.

[0066] Furthermore, between the first encoder and the decoder and between the second encoder and the decoder, a Channel Attention Cascade (CAC) module is added to strengthen the feature transfer between the first encoder and the decoder and the feature transfer between the second encoder and the decoder.

[0067] Specifically, referring to Figure 3, between the first encoder and the decoder, and between the second encoder and the decoder, a channel attention cascade module is used to replace the original skip connection. The so-called skip connection means adding the layers in the encoder to the layers of the same size in the decoder, and feature transfer between the encoder and the decoder is carried out through the skip connection. In this application, channel attention is used to change the weights on different channels in the two encoders, and then added together, which strengthens the feature transfer between the first encoder and the decoder and the feature transfer between the second encoder and the decoder. The channel attention cascade module is as shown in Figure 6 where C represents the number of channels, W represents the width, and H represents the length.

[0068] Therefore, in this embodiment, the channel attention cascade module is used to further improve the quality of video frame interpolation.

[0069] Step S300: Convolve the decoded features to obtain convolution parameters;

[0070] Specifically, multiple groups of standard convolutions are used to perform convolution operations on the decoded features to obtain convolution parameters. Referring to Figure 3 , the parameter W is obtained by performing convolution operations using a group of standard convolutions of (Conv + ReLU)×3, (Upsample + Conv + ReLU), and Softmax respectively. The parameter α is obtained by performing convolution operations using a group of standard convolutions of (Conv + ReLU)×3 and (Upsample + Conv + ReLU) respectively. The parameter β is obtained by performing convolution operations using a group of standard convolutions of (Conv + ReLU)×3 and (Upsample + Conv + ReLU) respectively. The parameter V is obtained by performing convolution operations using a group of standard convolutions of (Conv + ReLU)×3, (Upsample + Conv + ReLU), and Sigmoid respectively, where AvgPool represents average pooling, Softmax represents normalization, and Sigmoid represents the activation function.

[0071] Step S400: Based on the convolution parameters and two adjacent frames, obtain the target synthesized frame.

[0072] Based on the convolution parameters and two adjacent frames, calculate to obtain the target synthesized frame through the calculation formula. The calculation formula is:

[0073]

[0074] I t (i, j) = V⊙I n (i, j) + (1 - V)⊙I n+1 (i, j)(2)

[0075] In formula (1), I(i, j) represents the value of a pixel point obtained after the input frame is calculated by AdaCoF, K represents the size of the kernel, i represents the coordinate of the pixel point on the x-axis, j represents the coordinate of the pixel point on the y-axis, α k,l represents the vertical offset of the pixel point of the input frame, β k,l represents the horizontal offset of the pixel point of the input frame, W k,l represents the weight of the kernel, I(i + dk + α k,l , j + dl + β k,l ) represents the value of the pixel point of the input frame, dk represents the vertical dilation of the pixel point of the input frame, and dl represents the horizontal dilation of the pixel point of the input frame;

[0076] In formula (2), V represents the weight obtained by a group of convolutions. The value range of V is from 0 to 1. ⊙ represents element-wise multiplication. I n (i, j) represents the value of the pixel point obtained after the previous frame of two adjacent frames is calculated by formula (1), I n+1 (i, j) represents the value of the pixel point obtained after the subsequent frame of two adjacent frames is calculated by formula (1), I t (i, j) represents the value of the pixel point of the target synthesized frame.

[0077] Referring to Figure 7 , the embodiment of the present invention also provides a deep learning-based bidirectional encoding video frame interpolation system, including:

[0078] A fusion encoding feature acquisition unit 100, configured to acquire two adjacent frames in a target video, perform forward encoding on the two adjacent frames using a first encoder to obtain a first encoding feature, perform backward encoding on the two adjacent frames using a second encoder to obtain a second encoding feature, and fuse the first encoding feature and the second encoding feature to obtain a fusion encoding feature;

[0079] A decoded feature acquisition unit 200, configured to use a decoder to decode the fusion encoding feature to obtain a decoded feature;

[0080] A parameter acquisition unit 300, configured to perform convolution on the decoded feature to obtain convolution parameters;

[0081] A target synthesized frame acquisition unit 400, configured to calculate and obtain a target synthesized frame based on the convolution parameters and the two adjacent frames.

[0082] It should be noted that since the deep learning-based bidirectional encoding video frame interpolation system in this embodiment and the above-mentioned deep learning-based bidirectional encoding video frame interpolation method are based on the same inventive concept, the corresponding content in the method embodiment is equally applicable to the system embodiment of the present invention and will not be elaborated here.

[0083] An embodiment of the present invention also provides a deep learning-based bidirectional encoding video frame interpolation device, including: at least one control processor and a memory communicatively connected to the at least one control processor.

[0084] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0085] The non-transitory software programs and instructions required to implement the deep learning-based bidirectional encoding video frame interpolation method of the above embodiment are stored in the memory. When executed by the processor, the deep learning-based bidirectional encoding video frame interpolation method in the above embodiment is executed. For example, the method steps S100 to S400 described above are executed. Figure 1 in the method steps S100 to S400.

[0086] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0087] An embodiment of the present invention also provides a computer-readable storage medium storing computer-executable instructions, which are executed by one or more control processors, enabling the one or more control processors to execute the deep learning-based bidirectional encoding video frame interpolation method in the above method embodiment. For example, the functions of the method steps S100 to S400 described above are executed. Figure 1 in the method steps S100 to S400.

[0088] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform. Those skilled in the art can understand that all or part of the processes in the above-described method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the method embodiments as described above. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0089] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0090] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0091] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention. These equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.

Claims

1. A bidirectional encoding video frame interpolation method based on deep learning, characterized in that, It includes the following steps: Obtain two adjacent frames in the target video, perform forward encoding on the two adjacent frames using a first encoder to obtain first encoded features, perform backward encoding on the two adjacent frames using a second encoder to obtain second encoded features, and fuse the first encoded features and the second encoded features to obtain fused encoded features; Use a decoder to decode the fused encoded features to obtain decoded features; Perform convolution on the decoded features to obtain convolution parameters; Based on the convolution parameters and the two adjacent frames, calculate to obtain a target synthesized frame.

2. The method for bidirectional encoded video frame interpolation based on deep learning according to claim 1, wherein The performing forward encoding on the two adjacent frames using a first encoder to obtain first encoded features and performing backward encoding on the two adjacent frames using a second encoder to obtain second encoded features includes: Perform the forward encoding on the motion information of the two adjacent frames from the previous frame to the next frame using standard convolution of a low-dimensional feature layer and depthwise separable convolution of a high-dimensional feature layer in the first encoder to obtain the first encoded features; Perform the backward encoding on the motion information of the two adjacent frames from the next frame to the previous frame using standard convolution of a low-dimensional feature layer and depthwise separable convolution of a high-dimensional feature layer in the second encoder to obtain the second encoded features.

3. The bidirectional encoding video frame interpolation method based on deep learning according to claim 2, characterized in that The using a decoder to decode the fused encoded features to obtain decoded features includes: Perform decoding on the fused encoded features using standard convolution of a low-dimensional feature layer and depthwise separable convolution of a high-dimensional feature layer in the decoder to obtain the decoded features.

4. The bidirectional encoding video frame interpolation method based on deep learning according to claim 1, wherein It further includes: Add a channel attention cascade module between the first encoder and the decoder and between the second encoder and the decoder.

5. The method for bidirectional encoded video frame interpolation based on deep learning according to claim 1, characterized in that The performing convolution on the decoded features to obtain convolution parameters includes: Perform convolution operations on the decoded features using multiple groups of standard convolution to obtain the convolution parameters.

6. The bidirectional encoding video frame interpolation method based on deep learning according to claim 1, characterized in that The calculating to obtain a target synthesized frame based on the convolution parameters and the two adjacent frames includes: Based on the convolution parameters and the two adjacent frames, calculate to obtain a target synthesized frame through the following calculation formula, and the calculation formula is: I t (i, j) = V ⊙ I n (i, j) + (1 - V) ⊙ I n+1 (i, j) (2) In formula (1), I(i, j) represents the value of the pixel point obtained after the input frame is calculated by AdaCoF, K represents the size of the kernel, i represents the coordinate of the pixel point on the x-axis, j represents the coordinate of the pixel point on the y-axis, and α k,l represents the vertical offset of the pixel point of the input frame, and β k,l represents the horizontal offset of the pixel point of the input frame, W k,l represents the weight of the kernel, and I(i + dk + α k,l , j + dl + β k,l ) represents the value of the pixel point of the input frame, dk represents the vertical dilation of the pixel point of the input frame, and dl represents the horizontal dilation of the pixel point of the input frame; In formula (2), V represents the weights obtained by a set of convolutions. The value range of V is from 0 to 1. ⊙ represents pixel-by-pixel multiplication. I n (i, j) represents the value of the pixel point obtained by calculating the previous frame of two adjacent frames through formula (1). The I n+1 (i, j) represents the value of the pixel point obtained by calculating the subsequent frame of two adjacent frames through formula (1). The I t (i, j) represents the value of the pixel point of the target composite frame.

7. A bidirectional encoding video frame interpolation system based on deep learning, characterized in that, It includes: A fused encoded feature acquisition unit for obtaining two adjacent frames in the target video, performing forward encoding on the two adjacent frames using a first encoder to obtain first encoded features, performing backward encoding on the two adjacent frames using a second encoder to obtain second encoded features, and fusing the first encoded features and the second encoded features to obtain fused encoded features; A decoded feature acquisition unit for using a decoder to decode the fused encoded features to obtain decoded features; A parameter acquisition unit for performing convolution on the decoded features to obtain convolution parameters; A target synthesized frame acquisition unit for calculating to obtain a target synthesized frame based on the convolution parameters and the two adjacent frames.

8. A bidirectional encoding video frame interpolation device based on deep learning, characterized in that, Comprising at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the deep learning-based bidirectional encoded video frame interpolation method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the deep learning-based bidirectional encoded video frame interpolation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image encoding and decoding method and device for video sequence

    CN111263152A

  • Optical character recognition method and device, electronic equipment and storage medium

    CN113052156A