A low-latency video frame interpolation method and device based on space-time coding

CN115988162BActive Publication Date: 2026-09-25INNOSILICON MICROELECTRONICS (WUHAN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211538865.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-01
Publication Date
2026-09-25
Estimated Expiration
2042-12-01

AI Technical Summary

Technical Problem

然而,这些插帧方法中最快的对640P视频进行4x插帧所花时长也是原视频时长的2.6倍,可见相关插帧方法的实时性很弱、延时很高,难以满足对游戏、在线视频、直播等领域的落地应用

Benefits of technology

[0018]本发明实施例提供的基于时空编码的低延时视频插帧方法及设备,通过对输入进行时空编码,提取其特征后分别乘以待重建帧的时序编码矩阵,得到前后帧特征图对待重建帧有用的特征信息,摒弃了传统运动估计并位移补偿的思路,提升了计算速度并降低了延时,对于多输入共用特征提取块,节省了内存开销,提高了内存空间的利用效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115988162B_ABST
    Figure CN115988162B_ABST
Patent Text Reader

Abstract

The application provides a low-delay video frame interpolation method and device based on space-time coding. The method comprises the following steps: selecting a front frame image and a rear frame image from a time sequence section of a video to be interpolated, extracting corresponding front frame feature maps and rear frame feature maps from the front frame image and the rear frame image; activating the front frame feature maps and the rear frame feature maps, superimposing the activated front frame feature maps and the activated rear frame feature maps to obtain feature coding feature maps; decoding the feature coding feature maps to obtain a frame image to be inserted and construct an interpolation model; inputting a training set and a test set obtained in advance into the interpolation model for training and testing; and if the output of the interpolation model satisfies a preset threshold, determining that the interpolation model is a final interpolation model. The application reduces the delay and improves the utilization efficiency of the memory space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a low-latency video frame interpolation method and device based on spatiotemporal coding. Background Technology

[0002] Currently, video frame interpolation has undergone years of industry improvements, enhancing its performance. However, even the fastest interpolation method for 640P video (4x frame interpolation) still takes 2.6 times the original video length, demonstrating its weak real-time performance and high latency, making it unsuitable for practical applications in gaming, online video, and live streaming. Therefore, developing a low-latency video frame interpolation method and device to effectively overcome these shortcomings has become a pressing technical problem for the industry. Summary of the Invention

[0003] To address the aforementioned problems in the existing technology, embodiments of the present invention provide a low-latency video frame interpolation method and device based on spatiotemporal coding.

[0004] In a first aspect, embodiments of the present invention provide a low-latency video frame interpolation method based on spatiotemporal coding, comprising: selecting a previous frame image and a subsequent frame image from a temporal segment of a video to be interpolated; extracting corresponding previous frame feature maps and subsequent frame feature maps from the previous frame image and the subsequent frame image; activating the previous frame feature map and the subsequent frame feature map; superimposing the activated previous frame feature map and the activated subsequent frame feature map to obtain a feature-coded feature map; decoding the feature-coded feature map to obtain the frame image to be interpolated, and constructing an interpolation model; using a pre-acquired training set and test set as inputs to the interpolation model for training and testing; and determining the interpolation model as the final interpolation model if the output of the interpolation model meets a preset threshold.

[0005] Based on the above method embodiments, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention includes extracting corresponding previous frame feature maps and subsequent frame feature maps from the previous frame image and the subsequent frame image, comprising: synchronously inputting the previous frame image and the subsequent frame image into the feature extraction module to obtain the previous frame feature map and the subsequent frame feature map.

[0006] Based on the above method embodiments, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention includes activating the previous frame feature map and the next frame feature map, which includes: multiplying the previous frame feature map and the next frame feature map by the temporal coding matrix at the time of the frame to be interpolated, respectively, to obtain the encoded previous frame feature map and the encoded next frame feature map; and applying a nonlinear activation function to the encoded previous frame feature map and the encoded next frame feature map, respectively, to obtain the activated previous frame feature map and the activated next frame feature map.

[0007] Based on the above method embodiments, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention, wherein the step of superimposing the activated previous frame feature map and the activated subsequent frame feature map to obtain a feature-coded feature map includes: comparing one element of the image matrix of the activated previous frame feature map with another element at the corresponding position of the image matrix of the activated subsequent frame feature map, determining the smaller value between the one element and the other element as the superimposed element at the corresponding position, and superimposing all remaining elements of the image matrix of the activated previous frame feature map and the image matrix of the activated subsequent frame feature map in this manner to obtain a feature-coded feature map.

[0008] Based on the above method embodiments, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention includes decoding the feature-coded feature map to obtain the frame image to be inserted and constructing the frame interpolation model, which includes: inputting the feature-coded feature map into the feature reconstruction module, reconstructing the frame image to be inserted, and completing the construction of the frame interpolation model.

[0009] Based on the above method embodiments, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention includes the following feature extraction module: a spatiotemporal coding layer, which concatenates the encoding of each pixel in the duration, horizontal and vertical directions in the channel dimension of the input frame, and spatiotemporal coding is performed only once when the previous and next frames are input; a pyramid convolutional layer, which extracts features from the input feature map using convolutional kernels of different scales and concatenates them, picking features in different receptive fields; a first residual structure layer, which introduces the output of all nodes before the current node to prevent gradient vanishing; and a first batch normalization layer, which performs batch normalization after each convolutional layer to improve the robustness of the frame interpolation model.

[0010] Based on the above method embodiments, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention includes a feature reconstruction module comprising: a convolutional layer, a second residual structure layer, and a second batch normalization layer. The feature encoding feature map is reconstructed by using multi-layer convolution of the second residual structure layer and batch normalization of the second batch normalization layer.

[0011] Based on the above method embodiments, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention includes the following steps for obtaining the training set and test set: obtaining a first frame rate image set from the video according to a preset step size, extracting several frames from the first frame rate image set to form a second frame rate image set, and dividing the second frame rate image set into a training set and a test set.

[0012] Secondly, embodiments of the present invention provide a low-latency video frame interpolation device based on spatiotemporal coding, comprising: a first main module, configured to select a preceding frame image and a following frame image from a temporal segment of a video to be interpolated, and extract corresponding preceding frame feature maps and following frame feature maps from the preceding frame image and the following frame image; a second main module, configured to activate the preceding frame feature map and the following frame feature map, and superimpose the activated preceding frame feature map and the activated following frame feature map to obtain a feature-coded feature map; a third main module, configured to decode the feature-coded feature map to obtain the frame image to be interpolated, and construct an interpolation model; and a fourth main module, configured to train and test the interpolation model using a pre-acquired training set and test set as input, and if the output of the interpolation model meets a preset threshold, then the interpolation model is determined to be the final interpolation model.

[0013] Thirdly, embodiments of the present invention provide an electronic device, comprising:

[0014] At least one processor; and

[0015] At least one memory communicatively connected to the processor, wherein:

[0016] The memory stores program instructions that can be executed by the processor. The processor can call the program instructions to execute the low-latency video interpolation method based on spatiotemporal coding provided by any of the various implementations of the first aspect.

[0017] Fourthly, embodiments of the present invention provide a non-transitory computer-readable storage medium storing computer instructions that cause a computer to execute a low-latency video interpolation method based on spatiotemporal coding provided by any of the various implementations of the first aspect.

[0018] The low-latency video frame interpolation method and device based on spatiotemporal coding provided in this invention performs spatiotemporal coding on the input, extracts its features, and multiplies them by the temporal coding matrix of the frame to be reconstructed to obtain the feature maps of the preceding and following frames, which contain useful feature information for the frame to be reconstructed. This method abandons the traditional approach of motion estimation and displacement compensation, improves the calculation speed and reduces latency. For multiple inputs sharing the feature extraction block, it saves memory overhead and improves the efficiency of memory space utilization. Attached Figure Description

[0019] Figure 1 A flowchart of a low-latency video frame interpolation method based on spatiotemporal coding provided in an embodiment of the present invention;

[0020] Figure 2 A schematic diagram of the structure of a low-latency video frame interpolation device based on spatiotemporal coding provided in an embodiment of the present invention;

[0021] Figure 3A schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention;

[0022] Figure 4 This is a schematic diagram of the frame interpolation model structure provided in an embodiment of the present invention;

[0023] Figure 5 This is a schematic diagram illustrating the effect of reducing latency after video frame interpolation in an embodiment of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. In addition, the technical features of the various embodiments or individual embodiments provided by the present invention can be arbitrarily combined with each other to form feasible technical solutions. Such combinations are not constrained by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0025] This invention provides a low-latency video frame interpolation method based on spatiotemporal coding, see [link to relevant documentation]. Figure 1 The method includes: selecting a previous frame image and a subsequent frame image from the time segment of the video to be interpolated; extracting corresponding previous frame feature maps and subsequent frame feature maps from the previous frame image and subsequent frame image; activating the previous frame feature map and subsequent frame feature map; superimposing the activated previous frame feature map and activated subsequent frame feature map to obtain a feature-encoded feature map; decoding the feature-encoded feature map to obtain the frame image to be interpolated, and constructing an interpolation model; using a pre-acquired training set and test set as input for training and testing of the interpolation model; and determining the interpolation model as the final interpolation model if the output of the interpolation model meets a preset threshold.

[0026] In another embodiment, the task scene video (i.e., the video to be interpolated) is saved frame by frame as a high frame rate dataset. A low frame rate dataset is obtained by extracting one frame every n frames using a frame extraction method. This low frame rate dataset is then conventionally divided into a 7:3 ratio (or any other arbitrary ratio) to form a training set and a test set. Obtaining the low frame rate dataset from the high frame rate dataset using frame extraction is part of the dataset creation process. Deriving the high frame rate dataset from the low frame rate dataset is the result of the interpolation model calculation, commonly used for model testing.

[0027] Based on the above method embodiments, as an optional embodiment, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention includes extracting corresponding previous frame feature maps and subsequent frame feature maps from previous frame images and subsequent frame images, comprising: synchronously inputting the previous frame image and subsequent frame image into the feature extraction module to obtain the previous frame feature map and subsequent frame feature map.

[0028] In another embodiment, the preceding and following frames in the time sequence of the frames to be interpolated are selected as network inputs, and the preceding frame image is defined as F. t The image of the next frame is F. t+1 The image to be interpolated is F t+x Where t represents the time corresponding to the previous frame, x represents the delay between the frame to be inserted and the previous frame, and n represents the interval between any two frames in a low frame rate video; when there is only one frame before and after, n = 1. Furthermore, σ is defined as a non-linear activation function, ∩ is the intersection of matrix elements (the mathematical representation of a video image is a matrix), and the temporal coding matrix at time t+x is M. t+x The definition of the intersection ∩ of matrix elements is as follows: Assume but Where min represents the minimum value sign, a 11 a 12 a 21 and a 22 Describes the elements of matrix A, b 11 b 12 b 21 and b 22 Represents the elements of matrix B.

[0029] Based on the above method embodiments, as an optional embodiment, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention includes activating the previous frame feature map and the next frame feature map, which includes: multiplying the previous frame feature map and the next frame feature map by the temporal coding matrix at the time of the frame to be interpolated, respectively, to obtain the encoded previous frame feature map and the encoded next frame feature map; and using a nonlinear activation function to perform nonlinear activation on the encoded previous frame feature map and the encoded next frame feature map, respectively, to obtain the activated previous frame feature map and the activated next frame feature map.

[0030] In another embodiment, the frame interpolation model is constructed in two parts: encoding and decoding. In the encoding part, the previous frame image F... t and the following frame image F t+n The corresponding image feature map F′ is obtained by synchronously inputting the feature extraction block. t and F′ t+n The feature maps extracted from the preceding and following frames are combined with the temporal coding matrix M of the frame to be inserted. t+xThe two nonlinear activations are multiplied and then nonlinearly activated using the nonlinear activation function σ. The two nonlinear activation results are then superimposed using ∩ to obtain the feature-encoded predicted feature map of the frame to be interpolated. During decoding, the feature-encoded predicted feature map is input into the feature reconstruction module to reconstruct the image data of the frame to be interpolated.

[0031] Based on the above method embodiments, as an optional embodiment, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention, wherein the activation of the previous frame feature map and the activation of the subsequent frame feature map are superimposed to obtain a feature-coded feature map, includes: comparing one element of the image matrix of the activated previous frame feature map with another element at the corresponding position of the image matrix of the activated subsequent frame feature map, determining the smaller value between the one element and the other element as the superimposed element at the corresponding position, and superimposing all remaining elements of the image matrix of the activated previous frame feature map and the image matrix of the activated subsequent frame feature map in this manner to obtain a feature-coded feature map.

[0032] In another embodiment, the image matrix of the activated previous frame feature map is: The image matrix of the activated later frame feature map is In this matrix, each element of matrices C and D represents a pixel in the image. The activated feature maps of the previous and subsequent frames are superimposed, meaning the elements at corresponding positions in matrices C and D are compared, and the element with the smaller value (i.e., pixel) is selected as the element at that position, thus obtaining the feature-encoded feature map. E is the matrix representation of the feature-encoded feature map.

[0033] Based on the above method embodiments, as an optional embodiment, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention includes decoding the feature-coded feature map to obtain the frame image to be inserted and constructing the frame interpolation model, which includes: inputting the feature-coded feature map into the feature reconstruction module, reconstructing the frame image to be inserted, and completing the construction of the frame interpolation model.

[0034] Based on the above method embodiments, as an optional embodiment, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention includes the following feature extraction module: a spatiotemporal coding layer, which concatenates the encoding of each pixel in the duration, horizontal and vertical directions in the channel dimension of the input frame, and spatiotemporal coding is performed only once when the previous and next frames are input; a pyramid convolutional layer, which extracts features from the input feature map using convolutional kernels of different scales and concatenates them, picking features in different receptive fields; a first residual structure layer, which introduces the output of all nodes before the current node to prevent gradient vanishing; and a first batch normalization layer, which performs batch normalization after each convolutional layer to improve the robustness of the frame interpolation model.

[0035] In another embodiment, the feature extraction module includes: a spatiotemporal coding layer, a pyramid convolutional layer, a residual structure layer, and a batch normalization layer. The spatiotemporal coding layer concatenates unique codes for each pixel in the duration, horizontal, and vertical directions along the channel dimension of the input frame. Assuming the width and height of the input frame are H and W respectively, and a pixel's coordinates are (i,j), its duration, horizontal, and vertical codes can be written as:

[0036]

[0037] Where x represents the delay between the frame to be interpolated and the previous frame, and n represents the interval between any two frames in a low frame rate video. Spatiotemporal coding is performed only once when the preceding and following frames are input. If the frame is taken as the exact middle frame between the preceding and following frames, then x = 1 / 2, and the width and height of the frame are both 300 pixels. Therefore, the spatiotemporal coding can be written as:

[0038]

[0039] In the pyramid convolutional layer, convolutional kernels of different sizes are used to extract features from the input feature map and then concatenated, thereby picking up features from different receptive fields. In another embodiment, convolutional kernels of sizes 9, 7, 5, 3, and 1 can be used as a set of pyramid convolutions, with the feature map edges padded by 4, 3, 2, 1, and 0 pixels respectively to ensure consistent width and height of the convolution result. A residual structure is used to incorporate the outputs of all nodes before the current node to prevent gradient vanishing. Batch normalization is applied after each convolutional layer to improve the model's robustness.

[0040] Based on the above method embodiments, as an optional embodiment, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention includes a feature reconstruction module comprising: a convolutional layer, a second residual structure layer, and a second batch normalization layer. The feature encoding feature map is reconstructed by using multi-layer convolution of the second residual structure layer and batch normalization of the second batch normalization layer.

[0041] In another embodiment, the feature reconstruction module includes convolutional layers, residual structure layers, and batch normalization layers. Multiple convolutional layers with residual structure layers and batch normalization are used to reconstruct the feature map predicted by the feature encoding. 3×3 convolutions are used, with the last convolutional layer having 3 kernels to ensure the output has three RGB channels.

[0042] Based on the above method embodiments, as an optional embodiment, the low-latency video frame interpolation method based on spatiotemporal coding provided in this embodiment of the invention includes the following steps for obtaining the training set and test set: obtaining a first frame rate image set from the video according to a preset step size, extracting several frames from the first frame rate image set to form a second frame rate image set, and dividing the second frame rate image set into a training set and a test set.

[0043] Specifically, after the frame interpolation model is built, it is trained using a training set. This can be achieved by using the backpropagation algorithm on the training set, iterating until the mean squared error (MSE) loss value L ≤ 10. -2 If the accuracy of the interpolation model output after inputting both the training and test sets is greater than or equal to 98%, then this interpolation model is retained as the final interpolation model. Therefore, the final interpolation model is a trained and tested model that can be applied in practice. After obtaining the final interpolation model, it is quantized and pruned using TensorRT, then deployed to a hardware platform, utilizing the GPU to achieve low-latency, high-speed video interpolation. TensorRT is a C++ library that facilitates high-performance inference on NVIDIA Graphics Processing Units (GPUs), designed to work complementaryly with training frameworks such as TensorFlow, Caffe, PyTorch, and MXNet, specifically dedicated to fast and efficient network inference on GPUs. The interpolation effect can be seen in [link to documentation / example]. Figure 5 Compared to low frame rate sequences, 2.5x interpolation sequences significantly reduce the latency of 3 frames. After interpolation using the spatiotemporal coding-based low-latency video interpolation method provided in this invention, the video display latency is significantly reduced, the video playback process is smoother, and the video playback quality is improved.

[0044] The low-latency video frame interpolation method based on spatiotemporal coding provided in this invention performs spatiotemporal coding on the input, extracts its features, and multiplies them by the temporal coding matrix of the frame to be reconstructed to obtain the feature maps of the preceding and following frames that contain useful feature information for the frame to be reconstructed. This method abandons the traditional approach of motion estimation and displacement compensation, improves the calculation speed and reduces latency. For multiple inputs sharing the feature extraction block, it saves memory overhead and improves the efficiency of memory space utilization.

[0045] The constructed frame interpolation model structure can be found in [reference]. Figure 4 Previous frame image F t and the following frame image F t+n The synchronous input is fed into the feature extraction module (including spatiotemporal coding, convolutional layers, residual structures, and batch normalization layers) to obtain the corresponding image feature map F′. t and F′ t+n The feature maps extracted from the preceding and following frames are multiplied by the temporal coding matrix of the frame to be interpolated, and a nonlinear activation function σ is applied for nonlinear activation. The two nonlinear activation results are then superimposed using a logical AND operation (&) to obtain the feature-encoded predicted feature map of the frame to be interpolated. During decoding, the feature-encoded predicted feature map is input into the feature reconstruction module (including convolutional layers, residual structures, and batch normalization layers) to reconstruct the image F of the frame to be interpolated. t+x Where x is the delay between the frame to be inserted and the previous frame, and n is the interval between every two frames in a low frame rate video.

[0046] The various embodiments of this invention are implemented through programmed processing using a device with processor functionality. Therefore, in practical engineering, the technical solutions and functions of the various embodiments of this invention can be encapsulated into various modules. Based on this reality, and building upon the above embodiments, this invention provides a low-latency video frame interpolation apparatus based on spatiotemporal coding, which is used to execute the low-latency video frame interpolation method based on spatiotemporal coding in the above method embodiments. See also... Figure 2 The device includes: a first main module for selecting a preceding frame image and a following frame image from a time segment of a video to be interpolated, and extracting corresponding preceding frame feature maps and following frame feature maps from the preceding frame image and following frame image; a second main module for activating the preceding frame feature map and the following frame feature map, and superimposing the activated preceding frame feature map and the activated following frame feature map to obtain a feature-encoded feature map; a third main module for decoding the feature-encoded feature map to obtain the frame image to be interpolated, and constructing an interpolation model; and a fourth main module for training and testing the interpolation model using a pre-acquired training set and test set as input, and determining the interpolation model as the final interpolation model if the output of the interpolation model meets a preset threshold.

[0047] The low-latency video frame interpolation device based on spatiotemporal coding provided in this embodiment of the invention employs... Figure 2 Several modules in the model extract features by performing spatiotemporal encoding on the input and then multiplying them by the temporal encoding matrix of the frame to be reconstructed. This yields feature maps of the preceding and following frames that provide useful feature information for the frame to be reconstructed. This approach abandons the traditional motion estimation and displacement compensation method, improves computation speed and reduces latency. For multiple inputs sharing the feature extraction block, it saves memory overhead and improves memory space utilization efficiency.

[0048] It should be noted that the apparatus in the device embodiments provided by the present invention can be used not only to implement the methods in the above method embodiments, but also to implement the methods in other method embodiments provided by the present invention. The difference lies only in the setting of corresponding functional modules. Its principle is basically the same as that of the above device embodiments provided by the present invention. As long as those skilled in the art, based on the above device embodiments and referring to the specific technical solutions in other method embodiments, obtain corresponding technical means and technical solutions composed of these technical means by combining technical features, and improve the apparatus in the above device embodiments while ensuring the practicality of the technical solutions, they can obtain corresponding device-type embodiments for implementing the methods in other method-type embodiments. For example:

[0049] Based on the above device embodiments, as an optional embodiment, the low-latency video frame interpolation device based on spatiotemporal coding provided in this embodiment of the invention further includes: a first sub-module, used to implement the extraction of corresponding previous frame feature maps and subsequent frame feature maps from the previous frame image and the subsequent frame image, including: the previous frame image and the subsequent frame image are synchronously input into the feature extraction module to obtain the previous frame feature map and the subsequent frame feature map.

[0050] Based on the above-described device embodiments, as an optional embodiment, the low-latency video frame interpolation device based on spatiotemporal coding provided in this embodiment of the invention further includes: a second submodule, used to implement the activation of the previous frame feature map and the next frame feature map, including: multiplying the previous frame feature map and the next frame feature map with the temporal coding matrix at the time of the frame to be interpolated, respectively, to obtain the encoded previous frame feature map and the encoded next frame feature map; and using a nonlinear activation function to perform nonlinear activation on the encoded previous frame feature map and the encoded next frame feature map, respectively, to obtain the activated previous frame feature map and the activated next frame feature map.

[0051] Based on the above-described device embodiments, as an optional embodiment, the low-latency video frame interpolation device based on spatiotemporal coding provided in this embodiment of the invention further includes: a third submodule, used to superimpose the activated previous frame feature map and the activated subsequent frame feature map to obtain a feature-coded feature map, including: comparing one element of the image matrix of the activated previous frame feature map with another element at the corresponding position of the image matrix of the activated subsequent frame feature map, determining the smaller value between the one element and the other element as the superimposed element at the corresponding position, and superimposing all remaining elements of the image matrix of the activated previous frame feature map and the image matrix of the activated subsequent frame feature map in this manner to obtain a feature-coded feature map.

[0052] Based on the above device embodiments, as an optional embodiment, the low-latency video frame interpolation device based on spatiotemporal coding provided in this embodiment of the invention further includes: a fourth sub-module, used to decode the feature-coded feature map to obtain the frame image to be inserted and construct the frame interpolation model, including: inputting the feature-coded feature map into the feature reconstruction module, reconstructing the frame image to be inserted, and completing the construction of the frame interpolation model.

[0053] Based on the above device embodiments, as an optional embodiment, the low-latency video frame interpolation device based on spatiotemporal coding provided in this embodiment of the invention further includes: a fifth sub-module, used to implement the feature extraction module, including: a spatiotemporal coding layer, which concatenates the encoding of each pixel in the duration, horizontal and vertical directions in the channel dimension of the input frame, and spatiotemporal coding is only performed once when the previous and next frames are input; a pyramid convolutional layer, which extracts features from the input feature map and concatenates them using convolutional kernels of different scales, and picks up features in different receptive fields; a first residual structure layer, which introduces the output of all nodes before the current node to prevent gradient vanishing; and a first batch normalization layer, which performs batch normalization after each convolutional layer to improve the robustness of the frame interpolation model.

[0054] Based on the above device embodiments, as an optional embodiment, the low-latency video frame interpolation device based on spatiotemporal coding provided in this embodiment of the invention further includes: a sixth sub-module, used to implement the feature reconstruction module, including: a convolutional layer, a second residual structure layer and a second batch normalization layer, using multi-layer convolution of the second residual structure layer and batch normalization using the second batch normalization layer to calculate and reconstruct the feature-encoded feature map.

[0055] Based on the above device embodiments, as an optional embodiment, the low-latency video frame interpolation device based on spatiotemporal coding provided in this embodiment of the invention further includes: a seventh sub-module, used to acquire the training set and test set, including: acquiring a first frame rate image set from the video according to a preset step size, extracting several frames from the first frame rate image set to form a second frame rate image set, and dividing the second frame rate image set into a training set and a test set.

[0056] The method in this embodiment of the invention is implemented using an electronic device; therefore, it is necessary to introduce the relevant electronic device. For this purpose, this embodiment of the invention provides an electronic device, such as... Figure 3 As shown, the electronic device includes at least one processor, a communications interface, at least one memory, and a communications bus, wherein the at least one processor, the communications interface, and the at least one memory communicate with each other via the communications bus. The at least one processor can invoke logical instructions stored in the at least one memory to execute all or part of the steps of the methods provided in the foregoing method embodiments.

[0057] Furthermore, when the logical instructions in at least one of the aforementioned memories can be implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various method embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0058] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0059] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0060] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Based on this understanding, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or sometimes in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0061] It should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A low-latency video frame interpolation method based on spatiotemporal coding, characterized in that, include: Select the previous frame image and the next frame image from the time sequence of the video to be interpolated, and extract the corresponding previous frame feature map and the next frame feature map from the previous frame image and the next frame feature map. The feature maps of the previous frame and the next frame are activated, and the activated feature maps of the previous frame and the next frame are superimposed to obtain the feature encoding feature map. Decode the feature map to obtain the frame image to be inserted, and construct the frame insertion model; The frame interpolation model is trained and tested using pre-acquired training and test sets. If the output of the frame interpolation model meets a preset threshold, the frame interpolation model is determined to be the final frame interpolation model. The activation of the previous frame feature map and the next frame feature map includes: The feature maps of the previous frame and the next frame are multiplied by the temporal coding matrix of the frame to be inserted to obtain the feature map of the previous frame and the feature map of the next frame. The feature maps of the previous frame and the next frame are then nonlinearly activated by a nonlinear activation function to obtain the activated feature maps of the previous frame and the next frame. The step of superimposing the activated previous frame feature map and the activated subsequent frame feature map to obtain the feature-encoded feature map includes: One element of the image matrix of the activated previous frame feature map is compared with another element at the corresponding position of the image matrix of the activated subsequent frame feature map. The smaller value between the one element and the other element is determined as the superimposed element at the corresponding position. In this way, all the remaining elements of the image matrix of the activated previous frame feature map and the image matrix of the activated subsequent frame feature map are superimposed to obtain the feature-encoded feature map.

2. The low-latency video frame interpolation method based on spatiotemporal coding according to claim 1, characterized in that, The step of extracting the corresponding feature maps of the previous frame and the next frame from the previous frame image and the next frame image includes: The previous frame image and the next frame image are synchronously input into the feature extraction module to obtain the previous frame feature map and the next frame feature map.

3. The low-latency video frame interpolation method based on spatiotemporal coding according to claim 1, characterized in that, The process of decoding the feature-encoded feature map to obtain the frame image to be inserted and constructing the frame insertion model includes: The feature-encoded feature map is input into the feature reconstruction module to reconstruct the frame image to be inserted, thus completing the construction of the frame interpolation model.

4. The low-latency video frame interpolation method based on spatiotemporal coding according to claim 2, characterized in that, The feature extraction module includes: a spatiotemporal coding layer, which concatenates the encoding of each pixel in the three directions of duration, horizontal and vertical along the channel dimension of the input frame, with spatiotemporal coding performed only once between consecutive input frames; a pyramid convolutional layer, which extracts and concatenates features from the input feature map using convolutional kernels of different scales, picking features in different receptive fields; a first residual structure layer, which introduces the outputs of all nodes before the current node to prevent gradient vanishing; and a first batch normalization layer, which performs batch normalization after each convolutional layer to improve the robustness of the frame interpolation model.

5. The low-latency video frame interpolation method based on spatiotemporal coding according to claim 3, characterized in that, The feature reconstruction module includes: a convolutional layer, a second residual structure layer, and a second batch normalization layer. The feature encoding feature map is reconstructed by using multi-layer convolution of the second residual structure layer and batch normalization of the second batch normalization layer.

6. The low-latency video frame interpolation method based on spatiotemporal coding according to claim 5, characterized in that, The acquisition of the training set and the test set includes: acquiring a first frame rate image set from the video according to a preset step size, extracting several frames from the first frame rate image set to form a second frame rate image set, and dividing the second frame rate image set into a training set and a test set.

7. A low-latency video frame interpolation device based on spatiotemporal coding, characterized in that, include: The first main module is used to select the previous frame image and the next frame image from the time segment of the video to be inserted, and extract the corresponding previous frame feature map and the next frame feature map from the previous frame image and the next frame image. The second main module is used to activate the feature maps of the previous frame and the next frame, and then superimpose the activated feature maps to obtain the feature-encoded feature map. The third main module is used to decode the feature-encoded feature map to obtain the frame image to be inserted and construct the frame interpolation model. The fourth main module is used to train and test the frame interpolation model using a pre-acquired training set and test set. If the output of the frame interpolation model meets a preset threshold, the frame interpolation model is determined to be the final frame interpolation model. Activating the feature maps of the previous frame and the next frame includes: The feature maps of the previous frame and the next frame are multiplied by the temporal coding matrix of the frame to be inserted to obtain the feature map of the previous frame and the feature map of the next frame. The feature maps of the previous frame and the next frame are then nonlinearly activated by a nonlinear activation function to obtain the activated feature maps of the previous frame and the next frame. The activated feature map of the previous frame and the activated feature map of the subsequent frame are superimposed to obtain the feature-encoded feature map, which includes: One element of the image matrix of the activated previous frame feature map is compared with another element at the corresponding position of the image matrix of the activated subsequent frame feature map. The smaller value between the one element and the other element is determined as the superimposed element at the corresponding position. In this way, all the remaining elements of the image matrix of the activated previous frame feature map and the image matrix of the activated subsequent frame feature map are superimposed to obtain the feature-encoded feature map.

8. An electronic device, characterized in that, include: At least one processor, at least one memory, and a communication interface; wherein, The processor, memory, and communication interface communicate with each other; The memory stores program instructions that can be executed by the processor, which invokes the program instructions to perform the method of any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions that cause the computer to perform the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Video frame insertion method and device

    CN114821381A

  • Self-adaptive frame interpolation method based on space-time attention mechanism

    CN115147610A