Motion video frame rate enhancement method and system based on optical flow estimation and foreground detection
By combining foreground detection and optical flow estimation, using GAN to generate video complement frames, the problems of complex preprocessing and complex models in the prior art are solved, and more realistic video complement frames and frame rate enhancement effects are achieved.
Patent Information
- Application Number
- CN202111385904.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-22
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-11-22
AI Technical Summary
The existing video frame interpolation algorithm has complex preprocessing and complex models, making it difficult to generate realistic motion video complementary frames.
Combining foreground detection and optical flow estimation of moving targets, a video complement frame is generated using a generative adversarial network (GAN) to simulate the motion characteristics of the video moving targets.
Simplify preprocessing steps through foreground detection and optical flow estimation, reduce model complexity, make complement frames more realistic, and improve video frame rate and visual quality.
Smart Images

Figure CN114066761B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field related to computer vision, and in particular, to a method and system for enhancing the frame rate of a motion video based on optical flow estimation and foreground detection. Background Art
[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.
[0003] Due to the rapid development of multimedia technology, a variety of video sources with different frame rates have appeared on the market, and it is inevitable to convert frame rates between the above video sources. Frame rate up-conversion is a technology that converts low frame rate video into higher frame rate video. At the same time, as people's requirements for video quality continue to increase, high-definition video applications are becoming more and more common. Therefore, it is particularly important to study frame rate up-conversion algorithms suitable for high-definition video.
[0004] Video frame interpolation has a wide range of applications in computer vision, such as slow motion video, novel view synthesis, frame rate upconversion, and frame recovery in video streams. Videos with high frame rates can avoid common artifacts such as temporal jitter and motion blur, and are therefore more visually appealing to viewers. Video frame interpolation aims to synthesize intermediate frames between two consecutive video frames, which can be used to increase frame rates and enhance visual quality. Video frame interpolation is challenging due to the complex, large nonlinear motion and lighting changes in the real world. Currently, the commonly used methods for video frame interpolation algorithms in the prior art are: 1) warping the input frames according to the approximate optical flow; 2) using a convolutional neural network (CNN) to fuse and refine the warped frames. This method has problems such as complex preprocessing and complex models when performing video frame interpolation. Summary of the invention
[0005] In order to solve the problems of complex preprocessing and complex models in current methods, the present invention proposes a motion video frame rate enhancement method and system based on optical flow estimation and foreground detection. This method combines foreground detection and optical flow estimation of moving targets with the foreground motion information of the original video, and uses a generative adversarial network (GAN for short) to generate realistic motion video interpolation frames, which can simulate the motion characteristics of moving targets in the video and make the interpolation frames more realistic.
[0006] In order to achieve the above objectives, the present disclosure adopts the following technical solutions:
[0007] One or more embodiments provide a method for enhancing the frame rate of a motion video based on optical flow estimation and foreground detection, comprising the following steps:
[0008] The acquired video data to be processed is processed using a foreground detection algorithm to obtain a complete moving target image;
[0009] Use the optical flow estimation network to fit the video data to be processed to obtain the foreground optical flow information;
[0010] According to the target image and foreground optical flow information obtained by foreground detection, a generative adversarial network is used to generate video interpolation frames, and the video data to be processed is interpolated according to the video interpolation frames.
[0011] One or more embodiments provide a motion video frame rate enhancement system based on optical flow estimation and foreground detection, including a video acquisition device and a server;
[0012] Wherein, the video acquisition device is configured to acquire the video data to be enhanced and transmit it to the server;
[0013] The server is configured to execute the steps described in the above method.
[0014] One or more embodiments provide a motion video frame rate enhancement system based on optical flow estimation and foreground detection, including:
[0015] Foreground detection module: configured to process the acquired video data to be processed using a foreground detection algorithm to obtain a complete moving target image;
[0016] Optical flow estimation module: configured to obtain foreground optical flow information by fitting the video data to be processed using an optical flow estimation network;
[0017] The frame interpolation module is configured to generate video interpolation frames using a generative adversarial network based on the target image and foreground optical flow information obtained by foreground detection, and interpolate the original video data based on the video interpolation frames.
[0018] An electronic device comprises a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps described in the above method are completed.
[0019] Compared with the prior art, the present invention has the following beneficial effects:
[0020] The present disclosure combines foreground detection and optical flow estimation of moving targets with foreground motion information of the original video, and uses a generative adversarial network (GAN) to generate realistic motion video interpolation frames, which can simulate the motion characteristics of moving targets in the video and make the interpolation frames more realistic. In addition, through the early foreground detection and optical flow estimation of moving targets, the GAN model focuses on the interpolation operation of the moving target rather than the entire video frame, which is in line with the purpose of the interpolation operation, which is mainly to make the moving part smoother, and focuses on local interpolation rather than the whole, thereby greatly reducing the complexity of the model. At the same time, the present disclosure only performs foreground detection and optical flow estimation on the video frame, which simplifies the preprocessing steps.
[0021] Advantages of additional aspects of the present disclosure will be given in part in the following description and in part will become apparent from the following description or will be learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings constituting a part of the present disclosure are used to provide a further understanding of the present disclosure. The exemplary embodiments of the present disclosure and the description thereof are used to explain the present disclosure but do not constitute a limitation of the present disclosure.
[0023] Figure 1 is a schematic diagram of a motion video frame rate enhancement method according to Embodiment 1 of the present disclosure;
[0024] Figure 2 Schematic diagram of the structure of the generative adversarial network of Embodiment 1 of the present disclosure;
[0025] Figure 3 This is a flow chart of the motion video frame rate enhancement method according to Embodiment 1 of the present disclosure. DETAILED DESCRIPTION
[0026] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.
[0027] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present disclosure belongs.
[0028] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof. It should be noted that, in the absence of conflict, the various embodiments in the present disclosure and the features in the embodiments can be combined with each other. The embodiments will be described in detail below in conjunction with the accompanying drawings.
[0029] Example 1
[0030] In the technical solutions disclosed in one or more embodiments, Figure 1-3 As shown, the motion video frame rate enhancement method based on optical flow estimation and foreground detection includes the following steps:
[0031] Step S1, using a foreground detection algorithm to process the acquired video data to obtain a complete moving target image;
[0032] Step S2, using the optical flow estimation network to fit the video data to be processed to obtain foreground optical flow information;
[0033] Step S3: Generate video interpolation frames using a generative adversarial network based on the target image and foreground optical flow information obtained by foreground detection, and interpolate the original video data based on the video interpolation frames.
[0034] This embodiment combines foreground detection and optical flow estimation of moving targets to pre-process the input frame, and uses a generative adversarial network to generate realistic motion video interpolation frames, thereby achieving video frame rate enhancement. Through the early foreground detection and optical flow estimation of moving targets, the generative adversarial network focuses on the interpolation operation of the moving target rather than the entire video frame, which is consistent with the cognition that the interpolation operation is mainly to make the moving part smoother. Focusing on local interpolation rather than the overall greatly reduces the complexity of the model. Secondly, compared with the existing classic interpolation algorithms, such as the DAIN (Depth-Aware Video Frame Interpolation) algorithm, the preprocessing steps are simplified. In addition, using a generative adversarial network to generate video interpolation frames can simulate the motion characteristics of video moving targets, making the interpolation frames more realistic.
[0035] Furthermore, the method also includes the step of preprocessing the video data: converting the original video data and the enhanced high frame rate video true value into a frame sequence.
[0036] In this embodiment, the high frame rate video is relative to the original video image (mostly 30fps), and the frame rate of the video after frame supplementation is increased. A frame rate higher than the frame rate of the original video image is a high frame rate video.
[0037] Assume that the low frame rate original video is V in , the video frame rate is R in , assuming that the corresponding high frame rate video truth value after enhancement is V out , the frame rate is R out , each frame size in the video is w×h, the original video data and the high frame rate video true value are converted into the form of frame sequence to obtain the original image sequence Total N i Frame and Total N o frame.
[0038] The preprocessing of this embodiment converts the video into a frame sequence, and performs foreground detection and optical flow estimation on the input adjacent frames to meet the needs of the subsequent network frame supplementation operation. The preprocessing is simple and the processing speed is fast.
[0039] In step S1, a foreground detection algorithm is used to obtain a complete moving target image, and the specific method is as follows:
[0040] Step S11: for two adjacent frames of images in the video data to be processed, a background pixel distribution is modeled using a background modeling method to obtain a background image frame;
[0041] Step S12: Subtract the pixel values of the current video frame from the background image frame to obtain an image containing the complete moving target in the current video frame, thereby obtaining the moving target in the image.
[0042] In step S2, an optical flow estimation network is used to fit the video data to be processed to obtain foreground optical flow information. The optical flow estimation network used is PWC-Net, Pyramid, Warping, and Cost Volum, referred to as PWC.
[0043] In step S3, the structure of the Generative Adversarial Network (GAN) can be as follows: Figure 2 As shown, it includes a cascaded generator network and a discriminator network.
[0044] The generator network is used to generate realistic infill images, denoted as The number of interpolated frames N is determined by the original frame rate and the required generated video frame rate, that is, N = (R out / R in )-1;
[0045] The discriminator network is used to identify the interpolated frame image output by the generator network and the set video truth value V out If the error is within the set range, the interpolated frame generated by the generator is considered valid, otherwise it is invalid. The original video is interpolated according to the valid interpolated frame image to obtain an interpolated frame video with a high frame rate.
[0046] Among them, the interpolated frame image and the video truth value V out The error size can be determined by the loss function. The loss function is the sum of the absolute values of the differences between the pixel values of all generated video interpolation frames and the pixel values of the corresponding frames in the required enhanced high frame rate video. The formula is:
[0047]
[0048] Among them, c×w×h represents the number of all pixels, c represents the number of channels, w and h represent the width and height of the image, p represents each pixel in the image, and f i′ (p) represents the pixel value of the generated video frame, f j (p) represents the corresponding pixel value in the required enhanced high frame rate true value video.
[0049] Furthermore, the method further includes the step of training a generative adversarial network (GAN), including the following steps:
[0050] Step S31, obtaining original video data and performing preprocessing;
[0051] Step S32: Processing the acquired original video data using a foreground detection algorithm to obtain a complete moving target image;
[0052] Step S33, using the optical flow estimation network to fit the original video data to obtain foreground optical flow information;
[0053] Step S34, constructing a generative adversarial network, and generating video interpolation frames using the generative adversarial network according to the target image and foreground optical flow information obtained by foreground detection;
[0054] Step S35, constructing an objective function with the goal of minimizing the sum of the absolute values of the differences between the pixel values of the generated video interpolation frames and the corresponding pixel values in the required enhanced high frame rate video;
[0055] Step S36, using the back propagation algorithm and the stochastic gradient descent method to reduce the objective function to train the model, and optimizing and correcting the network weights based on the objective function, and obtaining the final video frame complement generation adversarial network after multiple iterative training.
[0056] The following is an explanation with a specific example.
[0057] Assume that the frame rate of the low frame rate original video is R in =60fps, assuming that the corresponding high frame rate video frame rate after frame addition is R out =120fps, and the size of each frame in the video is 256×256.
[0058] Video data preprocessing: converting the original video data and the high frame rate video true value into a frame sequence, obtaining a total of 1200 frames of the original video image sequence and a total of 2400 frames of the high frame rate video image sequence;
[0059] Step S1: Input two adjacent frames of images in sequence according to the time sequence, denoted as f i , f i+1 (i∈1,2…1199), using the foreground detection algorithm, that is, using the background modeling method to model the background pixel distribution, obtain the background image frame, subtract the pixel values of the current video frame from the background image frame to obtain the image containing the complete moving target in the current video frame, and obtain the moving target in the image;
[0060] Step S2: the two adjacent frames of images f input in step S1 i , f i+1 , use the optical flow estimation network fitting to calculate the foreground optical flow;
[0061] Combine the foreground target and foreground optical flow information obtained in steps S1 and S2, and use Figure 2The generator network in the GAN (Generative Adversarial Network) model shown in the figure generates a realistic interpolated image, denoted as f1, where the number of interpolated frames N = (R out / R in )-1=1.
[0062] Step S4, the generated video frame f1 and the corresponding high frame rate video true value are input into the discriminator of the GAN model to optimize the pixel-by-pixel L1 reconstruction loss function, that is, optimize the loss function value obtained by the loss function formula above; where c×w×h represents the number of all pixels, c represents the number of channels, such as the number of channels can be 3, w and h represent the width and height of the image, which is 256×256, p represents each pixel in the image, and f i′ (p) represents the pixel value of the generated video frame, f j (p) represents the corresponding pixel value in the high frame rate ground truth video.
[0063] Through the method of this embodiment, PWC (Pyramid, Warping, and Cost Volume)-Net is used as our optical flow estimation network, and the frame difference method is used as the foreground detection algorithm to qualitatively analyze the model complexity and the speed of the model generating interpolation frames. The model parameter amount of the current method is about 500,000, while the model parameter amount of the current classic video interpolation algorithm is at least about 1 million. Therefore, the proposed method greatly reduces the complexity of the model. In addition, the running time of the model to generate a frame of interpolation data is 8ms, and the running time of the current classic video interpolation algorithm is at least about 10ms, indicating that the method is simpler in preprocessing and has a shorter running time.
[0064] Example 2
[0065] Based on Embodiment 1, this embodiment provides a motion video frame rate enhancement system based on optical flow estimation and foreground detection, including a video acquisition device and a server;
[0066] Wherein, the video acquisition device is configured to acquire the video data to be enhanced and transmit it to the server;
[0067] The server is configured to execute the steps described in the method of Example 1.
[0068] The video acquisition device may be a camera.
[0069] Example 3
[0070] Based on Embodiment 1, this embodiment provides a motion video frame rate enhancement system based on optical flow estimation and foreground detection, including:
[0071] Foreground detection module: configured to process the acquired video data to be processed using a foreground detection algorithm to obtain a complete moving target image;
[0072] Optical flow estimation module: configured to obtain foreground optical flow information by fitting the video data to be processed using an optical flow estimation network;
[0073] The frame interpolation module is configured to generate video interpolation frames using a generative adversarial network based on the target image and foreground optical flow information obtained by foreground detection, and interpolate the original video data based on the video interpolation frames.
[0074] Example 4
[0075] Based on Example 1, this embodiment provides an electronic device, including a memory and a processor, and computer instructions stored in the memory and running on the processor. When the computer instructions are executed by the processor, the steps described in the method of Example 1 are completed.
[0076] The above description is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. For those skilled in the art, the present disclosure may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
[0077] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Technical personnel in the relevant field should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.
Claims
1. A motion video frame rate enhancement method based on optical flow estimation and foreground detection, characterized in that: The steps include: The acquired video data to be processed is processed using a foreground detection algorithm to obtain a complete moving target image; The foreground detection algorithm is used to obtain a complete moving target image. The specific method is as follows: For two adjacent frames of images in the video data to be processed, a background pixel distribution is modeled using a background modeling method to obtain a background image frame; Subtract the pixel values of the current video frame from the background image frame to obtain an image containing the complete moving target in the current video frame, and obtain the moving target in the image; Use the optical flow estimation network to fit the video data to be processed to obtain the foreground optical flow information; According to the target image and foreground optical flow information obtained by foreground detection, a generative adversarial network is used to generate video interpolation frames, and the video data to be processed is interpolated according to the video interpolation frames; The generative adversarial network consists of a cascaded generator network and a discriminator network; The generator network is used to generate realistic infill images; The discriminator network is used to identify the error between the interpolated frame image output by the generator network and the set video true value. When the error is within the set range, the original video is interpolated according to the interpolated frame image to obtain an enhanced interpolated frame video. The error between the interpolated frame image and the set video true value is determined by the loss function. The loss function is: the sum of the absolute values of the differences between the pixel values of all generated video interpolated frames and the pixel values of the corresponding frames in the required enhanced high frame rate video. The formula is: in, Represents the number of all pixels. Indicates the number of channels, and Indicates the width and height of the image. Represents each pixel in the image, Represents the pixel value of the generated video frame. Represents the corresponding pixel value in the required enhanced high frame rate true value video.
2. The motion video frame rate enhancement method based on optical flow estimation and foreground detection as claimed in claim 1, characterized in that: The method also includes the step of preprocessing the video data: converting the original video data and the high frame rate video truth value generated after enhancement into the form of a frame sequence.
3. The motion video frame rate enhancement method based on optical flow estimation and foreground detection as claimed in claim 1, characterized in that It also includes steps for training the generative adversarial network GAN, including the following: Get the original video data and perform preprocessing; The acquired original video data is processed by using the foreground detection algorithm to obtain a complete moving target image; Use the optical flow estimation network to fit the original video data to obtain the foreground optical flow information; Construct a generative adversarial network, and use it to generate video interpolation frames based on the target image and foreground optical flow information obtained by foreground detection; Construct the objective function of the generative adversarial network; The back propagation algorithm and stochastic gradient descent method are used to reduce the objective function to train the generative adversarial network model, and the network weights are optimized and corrected based on the objective function. After multiple iterative training, the final video frame complementation generative adversarial network is obtained.
4. The method for enhancing the frame rate of a moving video based on optical flow estimation and foreground detection as claimed in claim 3, characterized in that: The objective function of the generative adversarial network is to minimize the sum of the absolute values of the differences between the pixel values of the generated video interpolation frames and the corresponding pixel values in the required enhanced high frame rate video.
5. A motion video frame rate enhancement system based on optical flow estimation and foreground detection, characterized by: Including video acquisition device and server; Wherein, the video acquisition device is configured to acquire the video data to be enhanced and transmit it to the server; The server is configured to execute the steps described in any one of the methods of claims 1-4.
6. A motion video frame rate enhancement system based on optical flow estimation and foreground detection, characterized in that: include: Foreground detection module: configured to process the acquired video data to be processed using a foreground detection algorithm to obtain a complete moving target image; The foreground detection algorithm is used to obtain a complete moving target image. The specific method is as follows: For two adjacent frames of images in the video data to be processed, a background pixel distribution is modeled using a background modeling method to obtain a background image frame; Subtract the pixel values of the current video frame from the background image frame to obtain an image containing the complete moving target in the current video frame, and obtain the moving target in the image; Optical flow estimation module: configured to obtain foreground optical flow information by fitting the video data to be processed using an optical flow estimation network; A frame interpolation module is configured to generate video interpolation frames using a generative adversarial network according to the target image and foreground optical flow information obtained by foreground detection, and interpolate frames of the video data to be processed according to the video interpolation frames; The generative adversarial network consists of a cascaded generator network and a discriminator network; The generator network is used to generate realistic infill images; The discriminator network is used to identify the error between the interpolated frame image output by the generator network and the set video true value. When the error is within the set range, the original video is interpolated according to the interpolated frame image to obtain an enhanced interpolated frame video. The error between the interpolated frame image and the set video true value is determined by the loss function. The loss function is: the sum of the absolute values of the differences between the pixel values of all generated video interpolated frames and the pixel values of the corresponding frames in the required enhanced high frame rate video. The formula is: in, Represents the number of all pixels. Indicates the number of channels, and Indicates the width and height of the image. Represents each pixel in the image, Represents the pixel value of the generated video frame. Represents the corresponding pixel value in the required enhanced high frame rate true value video.
7. An electronic device, characterized in that: The method comprises a memory and a processor and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the steps described in any one of the methods of claims 1 to 4 are completed.
Citation Information
Patent Citations
Video generation method and device, storage medium and electronic equipment
CN110381268A
Video frame insertion method and device
CN112055249A