Video interpolation method and device, electronic equipment and medium
By combining the bidirectional gated loop unit with the Transformer module of the hierarchical multi-scale window mechanism, the performance limitations of the existing video interpolation model under the small sample data set are solved, efficient video frame interpolation is achieved, and the interpolation accuracy and robustness is improved. It is suitable for scenarios such as autonomous driving and game rendering.
Patent Information
- Application Number
- CN202510594116.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-09
AI Technical Summary
When the training data is insufficient, the generalization performance of the existing video interpolation model is weak, and the receptive field of traditional convolution operations is limited, resulting in limited video interpolation performance and it is difficult to achieve high-quality frame interpolation under small sample data sets.
The Transformer module adopts a bidirectional gating cyclic unit combined with a hierarchical multi-scale window mechanism to build a multi-scale spatiotemporal feature gating cyclic unit by integrating forward and reverse prediction results to improve the interpolation accuracy and robustness of the model under small sample data.
High-quality video frame interpolation is achieved under a small sample data set, improving the accuracy and consistency of interpolation frames, and is suitable for scenarios such as autonomous driving and game rendering.
Smart Images

Figure CN120128757A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and more particularly, to a video interpolation method, apparatus, electronic device, and medium. Background Art
[0002] The video interpolation task aims to repair or generate missing or damaged frames in a video through algorithms to restore the integrity and smoothness of the video, which is of great significance in the field of video processing. For example, the restoration of old movies and low-quality videos can significantly improve the overall quality and viewing experience of the film; the video interpolation technology in video surveillance systems can fill in the missing frames caused by network bandwidth limitations or device failures to ensure the integrity and continuity of surveillance videos; in the field of medical imaging, video interpolation can help generate more intermediate frames, improve the temporal resolution of dynamic images, enable doctors to more accurately observe the changes in the lesion area, and improve the accuracy of diagnosis, which is particularly important in dynamic medical imaging such as ultrasound scans.
[0003] Currently, the mainstream interpolation models are mainly based on the recurrent neural network (RNN) architecture. Since video data has temporal and spatial complexity, the model needs to capture the dependencies in both the temporal and spatial dimensions simultaneously. In recent years, researchers have proposed various hybrid models that combine convolutional neural networks (CNNs) and RNNs, such as ConvLSTM and E3D-LSTM, etc., to better model the spatio-temporal features in videos. However, traditional convolutional operations mainly focus on local spatial dependencies. Although the receptive field can be extended by stacking convolutional layers or increasing the convolutional kernel size, the actual effective receptive field is often much smaller than the theoretical value, which to some extent limits the video interpolation performance of the model. To solve this problem, the video vision transformer (ViViT) has been proposed, which models global spatio-temporal information through self-attention mechanisms, temporal encoding, and positional encoding. However, due to the lack of the inherent inductive bias ability of the Transformer architecture, its generalization performance is weak, and usually a large amount of training data is required to achieve ideal results. Based on this, there is a need to find a video frame interpolation method with less training data, faster speed, and high-quality generated images. Summary of the Invention
[0004] The purpose of the present invention is to provide a video interpolation method, apparatus, electronic device, and medium for the deficiencies of the prior art.
[0005] The purpose of the present invention is achieved through the following technical solutions: A video interpolation method includes:
[0006] Obtain the video to be interpolated and its missing segments, and construct an image data set with the video to be interpolated and its missing segments as a set of data;
[0007] Obtain the position information of the first frame of the missing segment in the video to be interpolated, and locate the interpolation range of the missing segment;
[0008] Use the image dataset to train the interpolation frame model. During the model training process, interpolate the video to be interpolated based on the position information of the first frame of the missing segment in the video to be interpolated and the interpolation range of the missing segment;
[0009] The trained interpolation frame model is used for video interpolation.
[0010] Furthermore, construct an image dataset with the video to be interpolated and its missing segment as a set of data, including:
[0011] Keep the one-to-one correspondence between the video to be interpolated and its missing segment;
[0012] The paired video data is decomposed frame by frame according to the time series, converted into a continuous sequence of image frames, and saved as image files;
[0013] Perform noise reduction processing on the images, then crop the images to the same resolution, and enhance the image data.
[0014] Furthermore, obtaining the position information of the first frame of the missing segment in the video to be interpolated and locating the interpolation range of the missing segment includes:
[0015] Obtain the image file of the first frame of the missing segment and use it as a reference frame; by calculating the structural similarity index between this reference frame and each frame in the image frame sequence of the video to be interpolated, determine the frame with the highest similarity, and use the next frame of this frame as the interpolation starting point, that is, the position of the first frame of the missing segment in the video to be interpolated;
[0016] Based on the interpolation starting point and the length of the missing segment, determine the interpolation range of the missing segment.
[0017] Furthermore, the interpolation frame model is a bidirectional gated recurrent unit.
[0018] The bidirectional gated recurrent unit includes a forward gated recurrent unit, a backward gated recurrent unit, and a two-stream feature fusion layer; the two-stream feature fusion layer uses a sum-average operation to fuse the image features of the bidirectional information flow;
[0019] Both the forward gated recurrent unit and the backward gated recurrent unit include image block encoding, a first multi-scale spatio-temporal feature gated recurrent unit, image block fusion, a second multi-scale spatio-temporal feature gated recurrent unit, image block expansion, and an image reconstruction layer, and are connected in sequence;
[0020] The image block encoding, image block fusion, and image block expansion are respectively used for image vectorization processing, feature dimensionality reduction aggregation, and feature upsampling; the image reconstruction layer uses two-dimensional transposed convolution operations to restore the spatial resolution and detail information of the image;
[0021] Both the first multi-scale spatio-temporal feature gated recurrent unit and the second multi-scale spatio-temporal feature gated recurrent unit include gated recurrent units.
[0022] Further, the gated recurrent unit nests a linear projection module and at least one Transformer module containing a hierarchical multi-scale window mechanism; the input of the gated recurrent unit is concatenated with the hidden state of the previous moment as the input of the linear projection module, and the output of the linear projection module is used as the input of the Transformer module containing the hierarchical multi-scale window mechanism; the output of the Transformer module containing the hierarchical multi-scale window mechanism is the output of the gated recurrent unit after passing through the key equation, which is also the output of the multi-scale spatio-temporal feature gated recurrent unit; the key equation is:
[0023] Update gate
[0024] Reset gate
[0025] Candidate hidden state
[0026] Hidden state
[0027] Among them, represents the Sigmoid activation function; represents the image frame input at the current moment; represents the hidden state of the previous moment; is the hyperbolic tangent function; , , , , and are all weight parameters of the model; , and are all biases of the model.
[0028] Further, if the gated recurrent unit nests multiple Transformer modules containing a hierarchical multi-scale window mechanism, the Transformer modules containing the hierarchical multi-scale window mechanism are connected in series.
[0029] Further, it also includes: only executing a forward gated recurrent unit for predicting the next image frame.
[0030] The present invention also provides a video interpolation device based on a bidirectional gated recurrent unit, including:
[0031] A dataset construction module, configured to obtain the video to be interpolated and its missing segments, and construct an image dataset with the video to be interpolated and its missing segments as a set of data;
[0032] An interpolation position and range acquisition module, configured to obtain the position information of the first frame of the missing segment in the video to be interpolated, and locate the interpolation range of the missing segment;
[0033] A model training module, configured to train an interpolation frame model using the image dataset. During the model training process, interpolate the video to be interpolated based on the position information of the first frame of the missing segment in the video to be interpolated and the interpolation range of the missing segment;
[0034] An interpolation module, configured to input the missing video to be processed into the trained interpolation frame model for interpolation, and output a complete video.
[0035] The present invention also provides an electronic device, including a memory and a processor, the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned video interpolation method based on a bidirectional gated recurrent unit.
[0036] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned video interpolation method based on a bidirectional gated recurrent unit is implemented.
[0037] Compared with the prior art, the beneficial effects of the embodiments of the present invention are:
[0038] 1. The bidirectional structure comprehensively captures the temporal dependencies in the video sequence by simultaneously modeling the historical and future information of the image frames, improving the accuracy and coherence of the interpolated frames. During the training process, the loss function is constrained by integrating the forward and backward prediction results, making full use of the bidirectional information flow to further improve the interpolation accuracy and robustness of the model.
[0039] 2. The multi-scale spatio-temporal feature gated recurrent unit, which combines the gated recurrent unit with the Transformer of the hierarchical multi-scale window mechanism, can efficiently extract the local details and global semantic features of the image under a small sample dataset, so as to achieve high-quality interpolation of the missing segments of the video.
[0040] 3. The method of the present invention can only execute the forward gated recurrent unit to predict the next image frame, and is applicable to scenarios such as an autonomous driving system to understand the surrounding dynamic environment and reduce game rendering latency. Description of the Drawings
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0042] Figure 1 Schematic flow chart of a video interpolation method based on a bidirectional gated recurrent unit provided by an embodiment of the present invention.
[0043] Figure 2 Schematic structural diagram of a bidirectional gated recurrent unit provided by an embodiment of the present invention.
[0044] Figure 3 Schematic structural diagram of a forward / backward gated recurrent unit provided by an embodiment of the present invention.
[0045] Figure 4 Schematic structural diagram of a multi-scale spatio-temporal feature gated recurrent unit provided by an embodiment of the present invention, where the activation function a is a Sigmoid activation function and the activation function b is a hyperbolic tangent function.
[0046] Figure 5 A schematic structural diagram of an electronic device corresponding to Figure 1 provided by the present invention. Detailed implementation manners
[0047] The following will describe the present invention in detail with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.
[0048] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0049] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms of "a", "the", and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0050] Figure 1 Schematic flow chart of a video interpolation method based on a bidirectional gated recurrent unit of the present invention, including the following steps:
[0051] S1, establish an image data set;
[0052] Specifically, the video to be interpolated and its missing segments must strictly maintain a one-to-one correspondence to ensure the accuracy and integrity of the data. The paired video data is decomposed frame by frame according to the time series, converted into a continuous sequence of image frames, and saved as image files; the Gaussian filter is used to denoise the images, and then the high-quality images are cropped to the same resolution, and the general data augmentation method in deep learning is used to augment the image data.
[0053] S2. Locate the interpolation range of the missing segment by obtaining the position information of the first frame of the missing segment in the video to be interpolated.
[0054] Specifically, obtain the first-frame image file of the missing segment and use it as the reference frame. By calculating the Structural Similarity Index (SSIM) between this reference frame and each frame in the sequence of image frames of the video to be interpolated, locate the frame with the highest similarity. Take the next frame of this frame as the interpolation starting point, and its subsequent frame sequence is the interpolation range of the missing segment, and the interpolation length is consistent with the length of the image frame sequence of the missing segment.
[0055] S3. Construct a bidirectional gated recurrent unit.
[0056] As Figure 2 shown, the bidirectional gated recurrent unit includes: a forward gated recurrent unit, a backward gated recurrent unit, and a two-stream feature fusion layer.
[0057] Figure 3 is the structural schematic diagram of the forward / backward gated recurrent unit, including the following structures: image patch encoding, two multi-scale spatio-temporal feature gated recurrent units, image patch fusion, image patch expansion, and image reconstruction layer; one multi-scale spatio-temporal feature gated recurrent unit is arranged between the image patch encoding and the image patch fusion, and the other is arranged between the image patch fusion and the image patch expansion.
[0058] The image patch encoding, image patch fusion, and image patch expansion are respectively the core technologies in the general Transformer architecture for image vectorization processing, feature dimensionality reduction aggregation, and feature upsampling. The image reconstruction layer uses two-dimensional deconvolution operation to restore the spatial resolution and detail information of the image.
[0059] The two-stream feature fusion layer uses an addition operation to fuse the image features of the bidirectional information flow.
[0060] As Figure 4 shown, the multi-scale spatio-temporal feature gated recurrent unit includes: a gated recurrent unit, a linear projection module, and a Transformer including a hierarchical multi-scale window mechanism.
[0061] The multi-scale spatio-temporal feature gated recurrent unit is an innovative architecture improved on the basis of the gated recurrent unit, that is, the gated recurrent unit is embedded with a linear projection module and a Transformer containing a hierarchical multi-scale window mechanism. One or more Transformers containing a hierarchical multi-scale window mechanism can be embedded. If multiple are embedded, they are connected in series, that is, the output of the first Transformer containing a hierarchical multi-scale window mechanism is the input of the second Transformer containing a hierarchical multi-scale window mechanism, and the output of the second Transformer containing a hierarchical multi-scale window mechanism is the input of the third Transformer containing a hierarchical multi-scale window mechanism, and so on; the hierarchical multi-scale window mechanism can capture local detail features and global semantic information simultaneously, thereby significantly enhancing the model's ability to model long-range dependencies. In the calculation process of the gated recurrent unit at each time step, it is mainly composed of the following four core components: (1) Update gate , which is used to control the update degree of the hidden state at the current moment and determine whether to integrate new information into the hidden state; (2) Reset gate , which is used to control the influence degree of the hidden state at the previous moment and determine whether to forget historical information; (3) Candidate hidden state , which generates new memory content based on the current input and the result of the reset gate as a potential candidate for state update; (4) Hidden state , which performs weighted fusion on the hidden state at the previous moment and the candidate hidden state through the update gate to determine the final state output, thereby realizing the dynamic update and retention of information. As Figure 3 shown, the key equation of the multi-scale spatio-temporal feature gated recurrent unit is:
[0062]
[0063]
[0064]
[0065]
[0066] Among them, represents the Sigmoid activation function; represents the image frame input at the current moment; represents the hidden state at the previous moment; is the hyperbolic tangent function, which compresses the result to [-1, 1]; , , , , and are all weight parameters of the model; , and are all biases of the model; is the Hadamard product operator.
[0067] The hierarchical multi-scale window mechanism decomposes the original global self-attention mechanism operation into local window calculations of multiple different scales, and performs feature extraction at different resolutions through a hierarchical processing method.
[0068] S4. The model evaluation uses the mean squared error (MSE), peak signal-to-noise ratio (PSNR), and structural similarity index (SSIM) as evaluation metrics to comprehensively measure the quality of the generated image frames and the similarity to the missing segments.
[0069] Example 1: A method for imputing missing segments based on ovarian cancer ultrasound images;
[0070] S1. Establish an image dataset;
[0071] Specifically, a high-resolution ultrasound device suitable for abdominal or pelvic scans is selected to scan the lesion site of ovarian cancer patients. By receiving and processing the echoes carrying the characteristic information of the lesion site, ultrasound images are obtained. Radiologists annotate the video segments containing the tumor site, and the data formats of the collected ultrasound and tumor site video segments are DICOM. The collected ultrasound images are the videos to be imputed, and the video segments of the tumor site are the missing segments, and the one-to-one correspondence between the videos to be imputed and their missing segments is maintained. The videos to be imputed and their missing segments are decomposed frame by frame according to the time sequence into image frame sequences, and the images are denoised using a Gaussian filter. After cropping to a resolution of 320 * 240, the data is enhanced using common deep learning data augmentation methods such as rotation, flipping, and scaling. In this example, the ovarian cancer ultrasound image dataset used contains a total of 68 ultrasound detection results. Each video to be imputed consists of 50 image frames, and the marked length of the missing segment is 10 frames. To ensure the training and evaluation effects of the model, the dataset is divided into a training set, a validation set, and a test set according to a ratio of 6:2:2, which are used for model training, hyperparameter tuning, and performance evaluation respectively.
[0072] S2. By obtaining the position information of the first frame of the missing segment in the video to be imputed, the imputation range of the missing segment is located;
[0073] Specifically, for each set of ultrasonic data, first, the first frame image of the missing segment is extracted as the reference frame. By calculating the structural similarity index between the reference frame and each frame in the sequence of video image frames to be interpolated, the frame with the highest similarity to the reference frame is located, and the index position of its next frame is recorded. During the model training process, this index is used as the interpolation starting point of the missing segment, and the missing content is predicted and generated according to a fixed interpolation length (10 frames).
[0074] S3. Construct a bidirectional gated recurrent unit;
[0075] Specifically, before and after the image block fusion operation, each forward gated recurrent unit and backward gated recurrent unit are respectively connected to two multi-scale spatio-temporal feature gated recurrent units to enhance the model's ability to extract local and global features.
[0076] Each multi-scale spatio-temporal feature gated recurrent unit integrates two groups of Transformer modules with a hierarchical multi-scale window mechanism (in this Embodiment 1, two groups of Transformer modules with a hierarchical multi-scale window mechanism are integrated, and the number of this module is not limited. It can also integrate 1 group, 3 groups or more than 3 groups). Therefore, each bidirectional gated recurrent unit contains a total of eight groups of Transformer modules with a hierarchical multi-scale window mechanism, which are respectively embedded into the forward and backward gated recurrent units. During the model training process, the batch size is set to 2, the Adam optimizer is used, the initial learning rate is set to 0.0001, and the mean squared error loss function is selected as the optimization target to minimize the pixel-level difference between the generated interpolated frame and the real frame;
[0077] S4. The model evaluation uses the mean squared error (MSE), peak signal-to-noise ratio (PSNR), and structural similarity index (SSIM) as evaluation metrics. For details, see Table 1.
[0078] Table 1: Evaluation Metrics
[0079] In this Embodiment 1, when the performance of the model reaches the convergence state on both the training set and the validation set, the complete network model structure and its corresponding hyperparameter configuration are saved. When calling this neural network algorithm in practical applications, the saved model file can be loaded, and the input data can be directly predicted according to the stored hyperparameters, so as to ensure the consistency of the model and the reliability of the prediction results.
[0080] This specification also provides a video interpolation device based on a bidirectional gated recurrent unit, including:
[0081] A dataset construction module, configured to obtain the video to be interpolated and its missing segments, and construct an image dataset with the video to be interpolated and its missing segments as a set of data;
[0082] An interpolation position and range acquisition module, configured to acquire the position information of the first frame of the missing segment in the video to be interpolated, and locate the interpolation range of the missing segment;
[0083] A model training module, configured to train an interpolation frame model using the image data set. During the model training process, interpolate the video to be interpolated based on the position information of the first frame of the missing segment in the video to be interpolated and the interpolation range of the missing segment;
[0084] An interpolation module, configured to input the missing video to be processed into the trained interpolation frame model for interpolation, and output a complete video.
[0085] It should be noted that the device embodiments shown in this embodiment match the content of the above method embodiments. The content of the above method embodiments can be referred to and will not be elaborated here.
[0086] This specification also provides Figure 5 a schematic structural diagram of an electronic device corresponding to Figure 1 as shown. As Figure 5 shown, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 video interpolation method.
[0087] This specification also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the above-mentioned video interpolation method based on a bidirectional gated recurrent unit.
[0088] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0089] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one or more flows and / or one or more blocks. Figure 1 one or more flows and / or one or more blocks Figure 1 a device for implementing the functions specified in one or more flows and / or one or more blocks.
[0090] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one or more flows and / or one or more blocks. Figure 1 one or more flows and / or one or more blocks Figure 1 a device for implementing the functions specified in one or more flows and / or one or more blocks.
[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, thereby providing steps for implementing the functions specified in one or more flows and / or one or more blocks by the instructions executed on the computer or other programmable device. Figure 1 one or more flows and / or one or more blocks Figure 1 a device for implementing the functions specified in one or more flows and / or one or more blocks.
[0092] The above embodiments are only used to illustrate the design concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.
Claims
1. A video interpolation method based on a bidirectional gated recurrent unit, characterized in that: include: Obtain the video to be interpolated and its missing segments, and construct an image data set using the video to be interpolated and its missing segments as a set of data; Obtaining the position information of the first frame of the missing segment in the video to be interpolated, and locating the interpolation range of the missing segment; Using the image data set to train an interpolation frame model, during the model training process, interpolating the video to be interpolated based on the position information of the first frame of the missing segment in the video to be interpolated and the interpolation range of the missing segment; The trained interpolation frame model is used for video interpolation.
2. A video interpolation method based on a bidirectional gated recurrent unit according to claim 1, characterized in that: An image dataset is constructed with the video to be interpolated and its missing segments as a set of data, including: Keep a one-to-one correspondence between the video to be interpolated and its missing segment; The paired video data is decomposed frame by frame according to the time series, converted into a continuous image frame sequence, and saved as an image file; The image is subjected to noise reduction, the image is cropped to the same resolution, and the image data is enhanced.
3. The video interpolation method based on bidirectional gated recurrent unit according to claim 1, characterized in that: Obtain the position information of the first frame of the missing segment in the video to be interpolated, and locate the interpolation range of the missing segment, including: The first frame image file of the missing segment is obtained and used as a reference frame; the frame with the highest similarity is determined by calculating the structural similarity index between the reference frame and each frame in the sequence of video image frames to be interpolated, and the next frame of the frame is used as the interpolation starting point, that is, the position of the first frame of the missing segment in the video to be interpolated; Based on the interpolation starting point and the length of the missing segment, the interpolation range of the missing segment is determined.
4. The video interpolation method based on bidirectional gated recurrent unit according to claim 1, characterized in that: The interpolation frame model is a bidirectional gated recurrent unit; The bidirectional gated recurrent unit includes a forward gated recurrent unit, a backward gated recurrent unit and a dual-stream feature fusion layer; the dual-stream feature fusion layer fuses the image features of the bidirectional information stream using a summing and averaging operation; The forward gated recurrent unit and the backward gated recurrent unit both include image block encoding, a first multi-scale spatiotemporal feature gated recurrent unit, image block fusion, a second multi-scale spatiotemporal feature gated recurrent unit, image block expansion and an image reconstruction layer, and are connected in sequence; The image block encoding, image block fusion and image block expansion are respectively used for image vectorization processing, feature dimension reduction aggregation and feature upsampling; the image reconstruction layer adopts a two-dimensional deconvolution operation to restore the spatial resolution and detail information of the image; The first multi-scale spatiotemporal feature gated recurrent unit and the second multi-scale spatiotemporal feature gated recurrent unit both include a gated recurrent unit.
5. The video interpolation method based on bidirectional gated recurrent unit according to claim 4, characterized in that: The gated recurrent unit is nested with a linear projection module and at least one Transformer module including a hierarchical multi-scale window mechanism; the input of the gated recurrent unit is concatenated with the hidden state of the previous moment as the input of the linear projection module, and the output of the linear projection module is used as the input of the Transformer module including the hierarchical multi-scale window mechanism; the output of the Transformer module including the hierarchical multi-scale window mechanism is the output of the gated recurrent unit after passing through the key equation; the key equation is: Update Gate Reset Gate Candidate hidden states Hidden State in, Represents the Sigmoid activation function; Represents the image frame input at the current moment; Indicates the hidden state at the previous moment; is the hyperbolic tangent function; , , , , and are all weight parameters of the model; , and are the biases of the model.
6. The video interpolation method based on bidirectional gated recurrent unit according to claim 5, characterized in that: If the gated recurrent unit nests multiple Transformer modules containing the hierarchical multi-scale window mechanism, the Transformer modules containing the hierarchical multi-scale window mechanism are connected in series.
7. A video interpolation method based on a bidirectional gated recurrent unit according to any one of claims 4 to 6, characterized in that: Also includes: Only the forward gated recurrent unit is executed to predict the next image frame.
8. A video interpolation device based on a bidirectional gated recurrent unit, characterized in that: include: A data set construction module is used to obtain the video to be interpolated and its missing segments, and construct an image data set using the video to be interpolated and its missing segments as a set of data; An interpolation position and range acquisition module is used to obtain the position information of the first frame of the missing segment in the video to be interpolated, and locate the interpolation range of the missing segment; A model training module, used to train an interpolation frame model using the image data set, and in the model training process, interpolating the video to be interpolated based on the position information of the first frame of the missing segment in the video to be interpolated and the interpolation range of the missing segment; The interpolation module is used to input the missing video to be processed into the trained interpolation frame model for interpolation and output the complete video.
9. An electronic device, comprising a memory and a processor, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the video interpolation method based on a bidirectional gated cyclic unit as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a video interpolation method based on a bidirectional gated recurrent unit as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and device for extracting key frame
CN111177460A
Video frame insertion method and device and server
CN112584232A
Video frame insertion method and device, equipment and storage medium
CN113014936A
Video shadow detection and elimination method based on deep learning
CN113378775A
Meteorological data interpolation method and device, electronic equipment and storage medium
CN113569972A
Cited By
Navigation semantic segmentation method based on hierarchical multi-scale Transform
CN120655928A