Video frame insertion method and device, computer equipment and computer storage medium
Through the collaboration of the optical flow feature extraction shared network and the storyboard neural network, combined with the optical flow interpolation module, the problem of distorted frames in video interpolation when the differences between shots are small is solved, and the video interpolation effect with high frame rate and smooth transition is achieved.
Patent Information
- Application Number
- CN202510916665.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-09-09
AI Technical Summary
Existing video interpolation algorithms based on optical flow are prone to produce distorted frames when the differences between shots are small, resulting in uneven transitions at the edges of the shots.
A multi-module collaborative algorithm framework consisting of a shared network for optical flow feature extraction, a storyboard neural network, and an optical flow interpolation module is adopted. Through optical flow feature extraction and guidance from storyboard results, the generation of distorted frames is avoided, achieving end-to-end high-frame-rate smooth video interpolation.
It effectively avoids distorted frames at the edge of the lens, generates interpolated videos with smooth transitions at the edge of the lens, and improves the smoothness of the video.
Smart Images

Figure CN120614477A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video frame insertion, and in particular to a video frame insertion method, device, computer equipment and computer storage medium. Background Art
[0002] Optical flow-based interpolation algorithms estimate the optical flow between two consecutive frames and use it to "deform" the original frames, achieving motion compensation and fusion, thereby generating intermediate frames and making the video smoother. However, this method can only address situations where the structural and color differences between shots are large and the differences within the shots are small. It still has significant limitations when the differences between shots are small. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to overcome the shortcomings of the prior art and provide a video interpolation method, device, computer equipment and computer storage medium, which are used to utilize a multi-module collaborative algorithm framework of an optical flow feature extraction shared network, a storyboard neural network and an optical flow interpolation module to automatically avoid "distorted frames" that may occur during interpolation and achieve end-to-end high-frame rate smooth video interpolation generation. It is necessary to uniformly extract the optical flow features of the video and guide the optical flow interpolation based on the storyboard results to avoid the occurrence of "distorted frames" and generate an interpolated video with a smooth transition at the edge of the lens.
[0004] The present invention provides the following technical solutions: In a first aspect, the present invention provides a video frame insertion method, comprising: Decompose the original video into multiple video frame images; Performing optical flow calculation on each of the video frame images through an optical flow feature extraction shared network to obtain multiple inter-frame optical flow features; Obtaining frame storyboard results corresponding to each of the video frame images according to each of the video frame images and each of the inter-frame optical flow features through a storyboard neural network; The optical flow interpolation module interpolates each of the video frame images according to the optical flow features between the frames and the frame segmentation results, so as to obtain an interpolated video after removing the distorted frames.
[0005] In one embodiment, the inter-frame optical flow features include a first forward optical flow feature between a previous frame video image and a current frame video image, a second forward optical flow feature between the current frame video image and a subsequent frame video image, and a first reverse optical flow feature between the subsequent frame video image and the current frame video image.
[0006] In one embodiment, obtaining the frame storyboard results corresponding to each of the video frame images according to each of the video frame images and each of the inter-frame optical flow features through the storyboard neural network includes: Extracting multi-dimensional features from each of the video frame images through the storyboard neural network; The frame segmentation results are obtained according to the multi-dimensional features, the first forward optical flow features and the second forward optical flow features.
[0007] In one embodiment, obtaining each frame segmentation result according to the multi-dimensional feature, the first forward optical flow feature, and the second forward optical flow feature includes: Obtaining a frame splitting probability corresponding to each of the video frame images according to the multi-dimensional feature, the first forward optical flow feature, and the second forward optical flow feature; Determine a binary result corresponding to each frame shot probability according to each frame shot probability and a preset probability threshold; Each of the binary results is used as each of the frame segmentation results.
[0008] In one embodiment, determining the binary result corresponding to each frame shot probability according to each frame shot probability and a preset probability threshold includes: If the frame splitting probability is greater than or equal to the preset probability threshold, the binary result corresponding to the frame splitting probability is a first value, and the first value is used to indicate that the frame video image corresponding to the frame splitting probability is a splitting point; If the frame splitting probability is less than the preset probability threshold, the binary result corresponding to the frame splitting probability is a second value, and the second value is used to indicate that the frame video image corresponding to the frame splitting probability is not a splitting point.
[0009] In one embodiment, the interpolating each of the video frame images according to the inter-frame optical flow features and the frame storyboard results by the optical flow interpolation module to obtain the interpolated video after removing the distorted frames includes: defining an intermediate frame video image between the current frame video image and the subsequent frame video image; For each of the current frame video images, obtaining, by the optical flow interpolation module, a third forward optical flow feature between the current frame video image and the intermediate frame video image, and a second reverse optical flow feature between the subsequent frame video image and the intermediate frame video image based on the second forward optical flow feature and the first reverse optical flow feature; Generating a fusion image according to the third forward optical flow feature and the second reverse optical flow feature, and obtaining a complement image of the fusion image; performing weighted fusion of the third forward optical flow feature and the second reverse optical flow feature according to the fusion image and the complement image to obtain a candidate intermediate frame; Determining whether the candidate intermediate frame is a shot edge intermediate frame according to the frame splitting results corresponding to the current frame video image and the subsequent frame video image; If the candidate intermediate frame is the shot edge intermediate frame, determining whether the shot edge intermediate frame is a distorted frame; If the shot edge intermediate frame is the distorted frame, generating a target intermediate frame according to the second reverse optical flow feature; The interpolated frame video is obtained according to the target intermediate frame and its corresponding current frame video image and subsequent frame video image.
[0010] In one embodiment, after determining whether the candidate intermediate frame is a shot edge intermediate frame, the method includes: If the candidate intermediate frame is not the shot edge intermediate frame, taking the candidate intermediate frame as the target intermediate frame; After determining whether the intermediate frame at the edge of the shot is a distorted frame according to the frame splitting results, the method includes: If the shot edge intermediate frame is not the warped frame, the candidate intermediate frame is used as the target intermediate frame.
[0011] In a second aspect, the present invention provides a video frame insertion device, comprising: A disassembly module, used to disassemble the original video into multiple video frame images; An extraction module is used to perform optical flow calculation on each of the video frame images through an optical flow feature extraction shared network to obtain multiple inter-frame optical flow features; A storyboard module is configured to obtain a frame storyboard result corresponding to each of the video frame images according to each of the video frame images and each of the inter-frame optical flow features through a storyboard neural network; The interpolation module is used to interpolate each of the video frame images according to the optical flow features between the frames and the frame segmentation results of each frame through the optical flow interpolation module to obtain an interpolated video with the distorted frames removed.
[0012] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the video frame insertion method as described in the first aspect is implemented.
[0013] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the video frame insertion method as described in the first aspect.
[0014] The video interpolation method, device, computer equipment and computer storage medium disclosed in the present invention decompose the original video into multiple video frame images; perform optical flow calculation on each of the video frame images through an optical flow feature extraction shared network to obtain multiple inter-frame optical flow features; obtain frame segmentation results corresponding to each of the video frame images based on each of the video frame images and each of the inter-frame optical flow features through a segmentation neural network; interpolate each of the video frame images based on each of the inter-frame optical flow features and each of the frame segmentation results through an optical flow interpolation module to obtain an interpolated video after removing distorted frames. In this way, through a multi-step and multi-round training framework, inter-frame optical flow features are obtained through an optical flow feature extraction shared network, the segmentation results provided by the segmentation neural network are used to guide the inspection and monitoring of distorted frames, the optical flow interpolation module is used to perform normal interpolation, and then the distorted part is replaced with a normal frame that only uses the optical flow of a single-side reference frame, thereby achieving end-to-end video interpolation generation and solving the problem of distorted frames at the edge of the lens. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope of protection of the present invention. In each of the drawings, similar components are numbered similarly.
[0016] Figure 1 A schematic diagram of a process of the video frame insertion method proposed in this embodiment is shown; Figure 2 Another schematic diagram of the process of the video frame insertion method proposed in this embodiment is shown; Figure 3 FIG2 shows another flow chart of the video frame insertion method proposed in this embodiment; Figure 4 A schematic diagram of the warped frame proposed in this embodiment is shown; Figure 5 Another schematic diagram of the process of the video frame insertion method proposed in this embodiment is shown; Figure 6 A structural diagram of the video frame insertion device proposed in this embodiment is shown.
[0017] Description of the accompanying drawings: 600 - video frame insertion device; 601 - disassembly module; 602 - extraction module; 603 - storyboard module; 604 - frame insertion module. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0019] The components of the embodiments of the present invention generally described and illustrated in the figures herein may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the figures is not intended to limit the scope of the claimed invention, but rather merely represents selected embodiments of the present invention. All other embodiments derived by those skilled in the art based on the embodiments of the present invention without inventive effort are intended to be within the scope of protection of the present invention.
[0020] Hereinafter, the terms "including", "having" and their cognates, which may be used in various embodiments of the present invention, are intended only to indicate specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be understood as first excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items or the possibility of adding one or more features, numbers, steps, operations, elements, components or combinations of the foregoing items.
[0021] Furthermore, the terms “first,” “second,” “third,” etc., are merely used for distinguishing descriptions and are not to be understood as indicating or implying relative importance.
[0022] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which the various embodiments of the present invention pertain. The terms (such as those defined in generally used dictionaries) will be interpreted as having the same meaning as in the context of the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning unless clearly defined in the various embodiments of the present invention.
[0023] Example 1 The disclosed embodiments provide a video interpolation method, which utilizes a multi-module collaborative algorithm framework of an optical flow feature extraction shared network, a storyboard neural network, and an optical flow interpolation module to automatically avoid "distorted frames" that may occur during interpolation and achieve end-to-end high-frame-rate smooth video interpolation generation. This method requires uniformly extracting the optical flow features of the video and guiding the optical flow interpolation based on the storyboard results to avoid the occurrence of "distorted frames" and generate an interpolated video with a smooth transition at the edge of the lens.
[0024] See Figure 1 The video frame insertion method includes steps S101 to S104, and each step is described in detail below.
[0025] Step S101: decompose the original video into multiple video frame images.
[0026] In this embodiment, the original video is decomposed into a plurality of continuous video frame images for frame insertion processing, which can improve the smoothness of the video after subsequent frame insertion.
[0027] Step S102 : performing optical flow calculation on each of the video frame images through an optical flow feature extraction shared network to obtain a plurality of inter-frame optical flow features.
[0028] In this embodiment, the shared network for optical flow feature extraction estimates inter-pixel motion by iteratively updating the optical flow field. It uses forward and backward consistency constraints to ensure optical flow accuracy and employs several constraint techniques to avoid unnatural mutations. The algorithm starts with a low-resolution image and gradually refines it to a high resolution, ultimately outputting an accurate optical flow estimate.
[0029] Specifically, see Figure 2 , extract the optical flow feature from the shared network and calculate the optical flow layer for the k-1th video frame image , the kth video frame image and the k+1th video frame image The optical flow features between frames are calculated to obtain multiple optical flow features between frames. With the current frame video image The first forward optical flow feature between , current frame video image and the next frame of video image The second forward optical flow feature between , and the next frame of video image With the current frame video image The first inverse optical flow feature between .
[0030] Step S103 , obtaining a frame storyboard result corresponding to each of the video frame images according to each of the video frame images and each of the inter-frame optical flow features through a storyboard neural network.
[0031] In this embodiment, after the optical flow features between each frame are spliced together through the storyboard neural network, they are projected as optical flow feature components to the storyboard feature projection layer in the storyboard neural network, thereby performing scene segmentation on the video image; at the same time, feature information is obtained from each video frame image through the storyboard neural network; further, the storyboard feature projection layer is combined with the feature information to determine whether each video frame image is a storyboard point based on the optical flow feature component, and this is used as the corresponding frame storyboard result to know the distorted frame recognition.
[0032] See Figure 3 In a specific embodiment, step S103 includes steps S1031 and S1032. Each step is described in detail below.
[0033] Step S1031: extracting multi-dimensional features from each of the video frame images through the storyboard neural network.
[0034] In this embodiment, a storyboard neural network is used to extract multi-dimensional features from each video frame image. The multi-dimensional features include structural features, color features, temporal convolution features, multi-dimensional features, etc.
[0035] Step S1032: Obtain the frame segmentation results according to the multi-dimensional features, the first forward optical flow features, and the second forward optical flow features.
[0036] In this embodiment, the feature projection layer collects multi-dimensional features, each first forward optical flow feature And each second forward optical flow feature Get the frame storyboard results corresponding to each video frame image.
[0037] In a specific embodiment, step S1032 includes: obtaining the frame splitting probability corresponding to each of the video frame images based on the multi-dimensional features, the first forward optical flow features, and the second forward optical flow features; determining the binary results corresponding to each of the frame splitting probabilities based on each of the frame splitting probabilities and a preset probability threshold; and using each of the binary results as each of the frame splitting results.
[0038] In this embodiment, the storyboard feature projection layer captures long-distance dependencies and performs feature enhancement based on multi-dimensional features, each first forward optical flow feature, and each second forward optical flow feature through an autocorrelation method, and finally obtains a more accurate frame storyboard probability; determines the binary results corresponding to each frame storyboard probability based on each frame storyboard probability and a preset probability threshold; and uses each binary result as a frame storyboard result.
[0039] Among them, the process of obtaining the binary result is: if the frame splitting probability is greater than or equal to the preset probability threshold, the binary result corresponding to the frame splitting probability is a first value, and the first value is used to indicate that the frame video image corresponding to the frame splitting probability is a splitting point, and the first value can be 1.
[0040] If the frame splitting probability is less than the preset probability threshold, the binary result corresponding to the frame splitting probability is a second value, which is used to indicate that the frame video image corresponding to the frame splitting probability is not a splitting point. The first value can be 0.
[0041] Step S104 , interpolating each of the video frame images according to the inter-frame optical flow features and the frame storyboard results by an optical flow interpolation module to obtain an interpolated video with the distorted frames removed.
[0042] In this embodiment, the optical flow interpolation module combines the optical flow features between frames to calculate the optical flow from the reference frame to the intermediate frame, thereby obtaining the intermediate frame. Furthermore, the frame splitting results are used to check for distortion in some important intermediate frames and remove possible distortions, obtaining the final interpolated video after removing the distorted frames. Finally, frame-by-frame merging is performed to obtain the final high-frame rate video. This implements video interpolation based on an end-to-end algorithm framework and three key modules. It can generate coherent interpolated frames for any type of video and any similarity of shot edges, avoiding the generation of distorted frames.
[0043] Please note that, see Figure 4 , the left and right are continuous video frames, and the middle is the distorted frame generated by optical flow interpolation. Through this embodiment, distorted frames between similar frames can be avoided and reasonable interpolation can be achieved.
[0044] See Figure 5 In a specific embodiment, step S104 includes steps S1041 to S1050, and each step is described in detail below.
[0045] Step S1041 : defining an intermediate frame video image between the current frame video image and the subsequent frame video image.
[0046] In this embodiment, the definition of the video image in the current frame and the next frame of video image Intermediate frames between video images , that is, the timestamp t is between k and k+1.
[0047] Step S1042: For each of the current frame video images, the optical flow interpolation module obtains a third forward optical flow feature between the current frame video image and the intermediate frame video image, and a second reverse optical flow feature between the subsequent frame video image and the intermediate frame video image based on the second forward optical flow feature and the first reverse optical flow feature.
[0048] In this embodiment, for each current frame video image, the optical flow interpolation module receives the second forward optical flow feature and the first reverse optical flow feature , use the inference model of deep learning method to get the current frame video image With intermediate frame video images The third forward optical flow feature between , and the next frame of video image With intermediate frame video images The second inverse optical flow feature between .
[0049] Step S1043: Generate a fusion image according to the third forward optical flow feature and the second reverse optical flow feature, and obtain a complement image of the fusion image.
[0050] In this embodiment, the third forward optical flow feature and the second reverse optical flow feature The image is aligned to the spatiotemporal position of the intermediate frame through deformable convolution or bilinear sampling, and is input into a lightweight convolutional network to extract multi-scale features and calculate the confidence or activity level of each pixel. Then, the network dynamically generates a fusion map M based on the aligned features, optical flow field and motion boundary information through the Softmax function, whose pixel values represent the third forward optical flow features. The weight contribution of Figure 1 − M It corresponds to the second reverse optical flow feature The weight of .
[0051] Step S1044 : performing weighted fusion on the third forward optical flow feature and the second reverse optical flow feature according to the fusion image and the complement image to obtain a candidate intermediate frame.
[0052] In this embodiment, according to the fusion graph M and the supplementary Figure 1 -M pairs of third forward optical flow features and the second reverse optical flow feature Perform weighted fusion to obtain candidate intermediate frames , then it is necessary to continue to perform distortion detection on the candidate intermediate frames. For example, .
[0053] Step S1045 , judging whether the candidate intermediate frame is a shot edge intermediate frame according to the frame storyboard results corresponding to the current frame video image and the subsequent frame video image.
[0054] In this embodiment, according to the current frame video image and the next frame of video image The corresponding frame storyboard result determines whether the candidate intermediate frame is a shot edge intermediate frame, so that the distortion and flickering problems existing in the transition part of the video shot can be avoided by identifying the shot edge intermediate frame.
[0055] It should be noted that for video frames a0, a1, …, ak, b0, b1, …, bk, where a and b each represent a shot, the frame between ak and the split point b0 is the shot edge intermediate frame. The split point result guides the inspection of that frame and nearby frames. If the previous split point result has already been obtained and a split point b0 is about to be generated, the intermediate frame before b0 is considered to need to be inspected.
[0056] Step S1046: If the candidate intermediate frame is the shot edge intermediate frame, determine whether the shot edge intermediate frame is a distorted frame.
[0057] In this embodiment, if the candidate intermediate frame is a shot edge intermediate frame, a distorted frame detection neural network is used to determine whether the shot edge intermediate frame is a distorted frame.
[0058] Step S1047: If the shot edge intermediate frame is the warped frame, a target intermediate frame is generated according to the second reverse optical flow feature.
[0059] In this embodiment, if the intermediate frame at the edge of the shot is a distorted frame, then according to the second reverse optical flow feature Generates a normal target intermediate frame. This replacement method will not cause too much impact and can remove the distortion effect.
[0060] Step S1048 , obtaining the interpolated frame video according to the target intermediate frame and its corresponding current frame video image and subsequent frame video image.
[0061] In this embodiment, each target intermediate frame and its corresponding current frame video image and the next frame of video image Merge and get the interpolated video after removing the distorted frames.
[0062] Step S1049: If the candidate intermediate frame is not the shot edge intermediate frame, use the candidate intermediate frame as the target intermediate frame.
[0063] In this embodiment, if the candidate intermediate frame is not a shot edge intermediate frame, the candidate intermediate frame is considered to be a normal frame and is directly used as the target intermediate frame.
[0064] Step S1050: If the shot edge intermediate frame is not the distorted frame, the candidate intermediate frame is used as the target intermediate frame.
[0065] In this embodiment, if the shot edge intermediate frame is not a distorted frame, then the shot edge intermediate frame is a normal frame, and the candidate intermediate frame can be directly used as the target intermediate frame.
[0066] The video interpolation method proposed in this embodiment decomposes the original video into multiple video frame images; performs optical flow calculation on each of the video frame images through an optical flow feature extraction shared network to obtain multiple inter-frame optical flow features; obtains frame segmentation results corresponding to each of the video frame images based on each of the video frame images and the inter-frame optical flow features through a segmentation neural network; and interpolates each of the video frame images based on the inter-frame optical flow features and the frame segmentation results through an optical flow interpolation module to obtain an interpolated video after removing distorted frames. In this way, through a multi-step, multi-round training framework, inter-frame optical flow features are obtained through an optical flow feature extraction shared network, the segmentation results provided by the segmentation neural network are used to guide distorted frame inspection and monitoring, the optical flow interpolation module is used to perform normal interpolation, and then the distorted part is replaced with a normal frame that only uses the optical flow of a single reference frame, thereby achieving end-to-end video interpolation generation and solving the problem of distorted frames at the edge of the lens.
[0067] Example 2 In addition, the present disclosure provides a video frame insertion device 600, see Figure 6 ,include: A decomposition module 601 is used to decompose the original video into multiple video frame images; An extraction module 602 is configured to perform optical flow calculation on each of the video frame images through an optical flow feature extraction shared network to obtain a plurality of inter-frame optical flow features; A storyboard module 603 is configured to obtain a frame storyboard result corresponding to each of the video frame images according to each of the video frame images and each of the inter-frame optical flow features through a storyboard neural network; The interpolation module 604 is configured to interpolate each of the video frame images according to the inter-frame optical flow features and the frame storyboard results through the optical flow interpolation module to obtain an interpolated video with the distorted frames removed.
[0068] Optionally, the inter-frame optical flow features include a first forward optical flow feature between a previous frame video image and a current frame video image, a second forward optical flow feature between the current frame video image and a subsequent frame video image, and a first reverse optical flow feature between the subsequent frame video image and the current frame video image.
[0069] Optionally, the storyboard module 603 is further used to extract multi-dimensional features based on each of the video frame images through the storyboard neural network; and obtain the storyboard results of each frame based on the multi-dimensional features, the first forward optical flow features and the second forward optical flow features.
[0070] Optionally, the splitting module 603 is further used to obtain the frame splitting probability corresponding to each of the video frame images based on the multi-dimensional features, the first forward optical flow features, and the second forward optical flow features; determine the binary results corresponding to each of the frame splitting probabilities based on each of the frame splitting probabilities and a preset probability threshold; and use each of the binary results as each of the frame splitting results.
[0071] Optionally, the mirroring module 603 is further configured to: if the frame mirroring probability is greater than or equal to the preset probability threshold, then the binary result corresponding to the frame mirroring probability is a first value, and the first value is used to indicate that the frame video image corresponding to the frame mirroring probability is a mirroring point; if the frame mirroring probability is less than the preset probability threshold, then the binary result corresponding to the frame mirroring probability is a second value, and the second value is used to indicate that the frame video image corresponding to the frame mirroring probability is not a mirroring point.
[0072] Optionally, the interpolation module 604 is further used to define an intermediate frame video image between the current frame video image and the subsequent frame video image; for each current frame video image, the optical flow interpolation module obtains a third forward optical flow feature between the current frame video image and the intermediate frame video image, and a second reverse optical flow feature between the subsequent frame video image and the intermediate frame video image according to the second forward optical flow feature and the first reverse optical flow feature; a fusion map is generated according to the third forward optical flow feature and the second reverse optical flow feature, and a complement map of the fusion map is obtained; according to the fusion The third forward optical flow feature and the second reverse optical flow feature are weightedly fused by the image and the complement image to obtain a candidate intermediate frame; whether the candidate intermediate frame is a shot edge intermediate frame is judged according to the frame splitting results corresponding to the current frame video image and the subsequent frame video image; if the candidate intermediate frame is the shot edge intermediate frame, whether the shot edge intermediate frame is a distorted frame is judged; if the shot edge intermediate frame is the distorted frame, a target intermediate frame is generated according to the second reverse optical flow feature; the interpolated frame video is obtained according to the target intermediate frame and its corresponding current frame video image and subsequent frame video image.
[0073] Optionally, the interpolation module 604 is further configured to use the candidate intermediate frame as the target intermediate frame if the candidate intermediate frame is not the shot edge intermediate frame; and use the candidate intermediate frame as the target intermediate frame if the shot edge intermediate frame is not the warped frame.
[0074] The device provided in the embodiment of the present disclosure can execute the steps of the video frame insertion method provided in Example 1, which will not be described again to avoid repetition.
[0075] The video interpolation device proposed in this embodiment breaks down the original video into multiple video frame images; performs optical flow calculation on each of the video frame images through an optical flow feature extraction shared network to obtain multiple inter-frame optical flow features; obtains frame segmentation results corresponding to each of the video frame images based on each of the video frame images and the inter-frame optical flow features through a segmentation neural network; and interpolates each of the video frame images based on the inter-frame optical flow features and the frame segmentation results through an optical flow interpolation module to obtain an interpolated video after removing distorted frames. In this way, through a multi-step, multi-round training framework, inter-frame optical flow features are obtained through an optical flow feature extraction shared network, the segmentation results provided by the segmentation neural network are used to guide distorted frame inspection and monitoring, the optical flow interpolation module is used to perform normal interpolation, and then the distorted part is replaced with a normal frame that only uses the optical flow of a single-side reference frame, thereby achieving end-to-end video interpolation generation and solving the problem of distorted frames at the edge of the lens.
[0076] Example 3 In addition, an embodiment of the present disclosure provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the video frame insertion method described in Example 1 is implemented.
[0077] The device provided in the embodiment of the present disclosure can execute the steps of the video frame insertion method provided in Example 1, which will not be described again to avoid repetition.
[0078] Example 4 The embodiment of the present disclosure provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the video frame insertion method described in the first embodiment is implemented.
[0079] In this embodiment, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0080] The computer-readable storage medium provided in this embodiment can implement the video frame insertion method provided in Example 1, and will not be described again here to avoid repetition.
[0081] In all examples shown and described herein, any specific values should be interpreted as merely exemplary and not limiting, and thus other examples of the exemplary embodiments may have different values.
[0082] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0083] The above-described embodiments merely illustrate several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, and all such variations and improvements fall within the scope of protection of the present invention.
Claims
1. A video frame insertion method, characterized in that: include: Decompose the original video into multiple video frame images; Performing optical flow calculation on each of the video frame images through an optical flow feature extraction shared network to obtain multiple inter-frame optical flow features; Obtaining frame storyboard results corresponding to each of the video frame images according to each of the video frame images and each of the inter-frame optical flow features through a storyboard neural network; The optical flow interpolation module interpolates each of the video frame images according to the optical flow features between the frames and the frame segmentation results, so as to obtain an interpolated video after removing the distorted frames.
2. The video frame insertion method according to claim 1, wherein: The inter-frame optical flow features include a first forward optical flow feature between a previous frame video image and a current frame video image, a second forward optical flow feature between the current frame video image and a subsequent frame video image, and a first reverse optical flow feature between the subsequent frame video image and the current frame video image.
3. The video frame insertion method according to claim 2, wherein: The obtaining of the frame storyboard results corresponding to each of the video frame images and each of the inter-frame optical flow features by using the storyboard neural network includes: Extracting multi-dimensional features from each of the video frame images through the storyboard neural network; The frame segmentation results are obtained according to the multi-dimensional features, the first forward optical flow features and the second forward optical flow features.
4. The video frame insertion method according to claim 3, wherein: The obtaining of each frame segmentation result according to the multi-dimensional feature, the first forward optical flow feature, and the second forward optical flow feature includes: Obtaining a frame splitting probability corresponding to each of the video frame images according to the multi-dimensional feature, the first forward optical flow feature, and the second forward optical flow feature; Determine a binary result corresponding to each frame shot probability according to each frame shot probability and a preset probability threshold; Each of the binary results is used as each of the frame segmentation results.
5. The video frame insertion method according to claim 4, wherein: Determining the binary result corresponding to each frame shot probability according to each frame shot probability and a preset probability threshold includes: If the frame splitting probability is greater than or equal to the preset probability threshold, the binary result corresponding to the frame splitting probability is a first value, and the first value is used to indicate that the frame video image corresponding to the frame splitting probability is a splitting point; If the frame splitting probability is less than the preset probability threshold, the binary result corresponding to the frame splitting probability is a second value, and the second value is used to indicate that the frame video image corresponding to the frame splitting probability is not a splitting point.
6. The video frame insertion method according to claim 2, wherein: The optical flow interpolation module interpolates each of the video frame images according to the inter-frame optical flow features and the frame storyboard results to obtain an interpolated video after removing the distorted frames, including: defining an intermediate frame video image between the current frame video image and the subsequent frame video image; For each of the current frame video images, obtaining, by the optical flow interpolation module, a third forward optical flow feature between the current frame video image and the intermediate frame video image, and a second reverse optical flow feature between the subsequent frame video image and the intermediate frame video image based on the second forward optical flow feature and the first reverse optical flow feature; Generating a fusion image according to the third forward optical flow feature and the second reverse optical flow feature, and obtaining a complement image of the fusion image; performing weighted fusion of the third forward optical flow feature and the second reverse optical flow feature according to the fusion image and the complement image to obtain a candidate intermediate frame; Determining whether the candidate intermediate frame is a shot edge intermediate frame according to the frame splitting results corresponding to the current frame video image and the subsequent frame video image; If the candidate intermediate frame is the shot edge intermediate frame, determining whether the shot edge intermediate frame is a distorted frame; If the shot edge intermediate frame is the distorted frame, generating a target intermediate frame according to the second reverse optical flow feature; The interpolated frame video is obtained according to the target intermediate frame and its corresponding current frame video image and subsequent frame video image.
7. The video frame insertion method according to claim 6, wherein: After determining whether the candidate intermediate frame is a shot edge intermediate frame, the method includes: If the candidate intermediate frame is not the shot edge intermediate frame, taking the candidate intermediate frame as the target intermediate frame; After determining whether the intermediate frame at the edge of the shot is a distorted frame according to the frame splitting results, the method includes: If the shot edge intermediate frame is not the warped frame, the candidate intermediate frame is used as the target intermediate frame.
8. A video frame insertion device, characterized in that: include: A disassembly module, used to disassemble the original video into multiple video frame images; An extraction module is used to perform optical flow calculation on each of the video frame images through an optical flow feature extraction shared network to obtain multiple inter-frame optical flow features; A storyboard module is configured to obtain a frame storyboard result corresponding to each of the video frame images according to each of the video frame images and each of the inter-frame optical flow features through a storyboard neural network; The interpolation module is used to interpolate each of the video frame images according to the optical flow features between the frames and the frame segmentation results of each frame through the optical flow interpolation module to obtain an interpolated video with the distorted frames removed.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the video frame insertion method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that It stores a computer program, which, when executed by a processor, implements the video frame insertion method according to any one of claims 1 to 7.