Event-based video frame interpolation method, device and equipment and storage medium
Patent Information
- Application Number
- CN202310400487.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-10
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-04-10
AI Technical Summary
[0005]本申请实施例提供了一种基于事件引导的视频插帧方法、装置、设备及存储介质,至少能够解决相关技术中难以保证卷帘曝光模式所采集的视频的插帧效果的问题
[0010] As can be seen from the above, according to the event-guided video frame interpolation method, apparatus, device, and storage medium provided in this application, for each adjacent shutter exposure image frame in the video to be interpolated, an event data stream collected within the exposure time interval of the adjacent shutter exposure image frames is acquired; the corresponding global deformation field is estimated based on the event data stream; a potential global exposure image frame combination corresponding to the adjacent shutter exposure image frames is generated based on the global deformation field; and different potential global exposure image frame combinations are inserted between the corresponding adjacent shutter exposure image frames to obtain the interpolated video. Through the implementation of this application, the conversion between shutter exposure image frames and global exposure image frames is realized based on the deformation field, and the prediction of potential image frames is guided by synchronously acquired event data. The event signal can provide effective support for the estimation of inter-frame motion state, thereby improving the frame interpolation effect of the video acquired in the shutter exposure mode.
Smart Images

Figure CN116489524B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more particularly to the field of computer vision technology, and can be applied to video frame interpolation scenarios. More specifically, this application discloses an event-guided video frame interpolation method, apparatus, device, and storage medium. Background Technology
[0002] Currently, users use terminal devices to shoot videos in many scenarios. As technology continues to develop, users have higher and higher requirements for video frame rates. However, ordinary cameras are limited by their configuration levels and cannot meet users' high frame rate needs. Based on this, video frame interpolation technology has emerged.
[0003] Video frame interpolation is a commonly used technique in video enhancement, aiming to convert low frame rate video into high frame rate video. In practical applications, rolling shutter exposure is the most common exposure mode, which often results in image and video distortion and warping. Currently, event-guided video frame interpolation techniques only consider global exposure modes and cannot meet the frame interpolation requirements of rolling shutter exposure. Furthermore, current techniques typically use raw APS image frames to predict potential image frames, but the guiding information provided by the raw APS image frames is limited, making it difficult to guarantee the accuracy of potential image frame prediction, resulting in poor video frame interpolation performance.
[0004] It is important to note that the techniques described in this section are not necessarily those previously conceived or adopted. Unless otherwise specified, no technique described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be recognized in any prior art. Summary of the Invention
[0005] This application provides an event-guided video frame interpolation method, apparatus, device, and storage medium, which can at least solve the problem in related technologies that it is difficult to guarantee the frame interpolation effect of videos acquired in the rolling shutter exposure mode.
[0006] The first aspect of this application provides an event-guided video frame interpolation method, comprising: acquiring an event data stream collected within the exposure time interval of each adjacent shutter exposure image frame in the video to be interpolated; estimating a corresponding global deformation field based on the event data stream; generating a potential global exposure image frame combination corresponding to the adjacent shutter exposure image frames according to the global deformation field; wherein the potential global exposure image frame combination includes multiple potential global exposure image frames; and inserting different potential global exposure image frame combinations between the corresponding adjacent shutter exposure image frames to obtain a video with frame interpolation completed.
[0007] A second aspect of this application provides an event-guided video frame interpolation apparatus, comprising: an acquisition module, configured to acquire an event data stream collected within the exposure time interval of each adjacent shutter exposure image frame in the video to be interpolated; an estimation module, configured to estimate a corresponding global deformation field based on the event data stream; a generation module, configured to generate a potential global exposure image frame combination corresponding to the adjacent shutter exposure image frames based on the global deformation field; wherein the potential global exposure image frame combination includes multiple potential global exposure image frames; and an interpolation module, configured to insert different potential global exposure image frame combinations into the corresponding adjacent shutter exposure image frames to obtain a frame-interpolated video.
[0008] A third aspect of this application provides an electronic device, including a memory and a processor, wherein the processor is configured to execute a computer program stored in the memory, and when the processor executes the computer program, it implements the steps of the video frame interpolation method provided in the first aspect of this application.
[0009] The fourth aspect of this application provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the steps of the video frame interpolation method provided in the first aspect of this application.
[0010] As can be seen from the above, according to the event-guided video frame interpolation method, apparatus, device, and storage medium provided in this application, for each adjacent shutter exposure image frame in the video to be interpolated, an event data stream collected within the exposure time interval of the adjacent shutter exposure image frames is acquired; the corresponding global deformation field is estimated based on the event data stream; a potential global exposure image frame combination corresponding to the adjacent shutter exposure image frames is generated based on the global deformation field; and different potential global exposure image frame combinations are inserted between the corresponding adjacent shutter exposure image frames to obtain the interpolated video. Through the implementation of this application, the conversion between shutter exposure image frames and global exposure image frames is realized based on the deformation field, and the prediction of potential image frames is guided by synchronously acquired event data. The event signal can provide effective support for the estimation of inter-frame motion state, thereby improving the frame interpolation effect of the video acquired in the shutter exposure mode.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0012] The accompanying drawings exemplify embodiments and form part of the specification, working together with the textual description to explain exemplary implementations of the embodiments. The drawings shown are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0013] Figure 1 This is a schematic diagram of the basic process of a video frame interpolation method provided in an embodiment of this application; Figure 2 A schematic diagram illustrating the association between an event data stream and adjacent shutter exposure image frames provided in an embodiment of this application; Figure 3 A schematic diagram illustrating the principle of obtaining an initialized global exposure image frame according to an embodiment of this application; Figure 4 This is a schematic diagram of a video frame interpolation effect provided in an embodiment of this application; Figure 5 A detailed flowchart illustrating a video frame interpolation method provided in an embodiment of this application; Figure 6 A schematic diagram of the program modules of a video frame interpolation device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0014] To make the inventive objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] It should be understood that in the description of the embodiments of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0016] The following describes in detail, with reference to the accompanying drawings, an event-guided video frame interpolation method, apparatus, device, and storage medium according to embodiments of this application.
[0017] To address the issue of ensuring effective frame interpolation in videos captured using the rolling shutter exposure mode in related technologies, one embodiment of this application provides an event-guided video frame interpolation method, such as... Figure 1 A basic flowchart of a video frame interpolation method provided in an embodiment of this application specifically includes the following steps: Step 101: For each adjacent shutter exposure image frame in the video to be interpolated, acquire the event data stream collected in the exposure time interval of the adjacent shutter exposure image frames.
[0018] Specifically, in this embodiment, the shutter exposure image frame can be acquired by an active-pixel sensor (APS), while the event data stream is acquired based on a dynamic vision sensor (DVS). In practical applications, the active-pixel sensor and the dynamic vision sensor in this embodiment can be discrete image sensors or integrated image sensors (or fusion sensors). The overall photosensitive area of the integrated image sensor can be divided into multiple sub-photosensitive areas, and the pixel arrays of the multiple sub-photosensitive areas correspond to the APS data mode and the DVS data mode, respectively. Compared with multiple discrete sensor modules, this effectively compresses the device size of the sensor module and is more conducive to the miniaturization of the overall hardware architecture.
[0019] It should be noted that the dynamic vision sensor is a new type of sensor that simulates the human retina and responds to pixel pulses caused by changes in brightness due to motion. Therefore, it can capture changes in scene brightness (i.e., changes in light intensity) at an extremely high frame rate, record events at specific points in time and at specific locations in the image, forming an event stream rather than a frame stream. This can solve the problems of information redundancy, large data storage and real-time processing requirements of traditional cameras.
[0020] In this embodiment, after each pixel of the dynamic vision sensor generates a photoelectric analog signal at the current working moment, the difference between its current photoelectric analog signal and the previous photoelectric analog signal is calculated. Then, the signal difference is compared with a preset difference threshold to obtain the event polarity of the corresponding pixel. Finally, event data is generated based on the pixel coordinates, the current timestamp, and the event polarity. The event data is represented as follows: ,in, Indicates the current working time. This represents the pixel coordinate position of a pixel. This indicates the event polarity. When the signal difference is greater than a preset first difference threshold, the event polarity is 1, indicating that the pixel has generated a positive event; when the signal difference is less than a preset second difference threshold, the event polarity is -1, indicating that the pixel has generated a negative event; when the signal difference is greater than or equal to the second difference threshold and less than or equal to the first difference threshold, the event polarity is 0, indicating that the pixel has not generated an event. Different event polarities are related to the comparison of signals generated by the same pixel at different times in a dynamic object. It is worth noting that the first difference threshold is positive and the second difference threshold is negative, and the two difference thresholds are preferably opposites of each other.
[0021] In one optional implementation of this embodiment, when acquiring event data streams collected within the exposure time interval of adjacent roller shutter exposure image frames, all event data streams between the first row exposure start time and the last row exposure end time of each roller shutter exposure image frame in the adjacent roller shutter exposure image frames can be acquired. The row exposure time interval of each row of the roller shutter exposure image frames is traversed, and redundant event data not within the row exposure time interval is deleted based on the timestamp of the event data stream, resulting in the remaining event data. All remaining event data corresponding to adjacent roller shutter exposure image frames are determined as the event data streams synchronously acquired within the overall exposure time interval. Based on this, event data not within the current row's exposure time is deleted according to the event data's timestamp, leaving only the event data truly corresponding to the adjacent roller shutter exposure image frames, effectively removing redundant event data.
[0022] like Figure 2 The diagram shown is a schematic representation of the association between an event data stream and adjacent shutter exposure image frames provided in this embodiment. and This represents two adjacent shutter exposure image frames, and Events represents the corresponding event signals. and It is The matrix, and Indicates the height and width of the image frame. This represents three color channels (R, G, B).
[0023] Step 102: Estimate the corresponding global deformation field based on the event data stream.
[0024] Specifically, in this embodiment, the event data stream can be divided into multiple time segments, and each time segment can be divided into multiple voxel frames to obtain the event matrix corresponding to the event data stream; the event matrix is traversed, and each adjacent time segment is input into the optical flow estimation network to perform optical flow estimation to obtain the optical flow corresponding to each adjacent time segment; the optical flow field composed of all optical flows is determined as the global deformation field (DF, Displacement Field) corresponding to the event data stream.
[0025] In this embodiment, the event data stream is first segmented into... Each time segment is divided into several time segments, and in this embodiment, each time segment is further divided into several time segments. Individual pixel frames, thereby transforming the event data stream into The event matrix is then used. Based on this, the event matrix is traversed along the event dimension, and adjacent time segments are input into the optical flow estimation network to obtain the optical flow within those adjacent time segments. This results in a dense optical flow field along the time dimension, with the shape of... In this embodiment, the optical flow field is used as the deformation field.
[0026] Step 103: Generate a combination of potential global exposure image frames corresponding to adjacent shutter exposure image frames based on the global deformation field.
[0027] Specifically, this embodiment uses a global deformation field generated based on event data stream to predict adjacent shutter exposure image frames. That is, it uses event data with high dynamic characteristics to guide the prediction of potential global exposure image frames of adjacent shutter exposure image frames. Event signals can greatly help estimate the motion state between frames, and finally predict a combination of potential global exposure image frames including multiple potential global exposure image frames.
[0028] In one optional implementation of this embodiment, firstly, adjacent shutter exposure image frames are transformed into initial global exposure image frames based on the global deformation field; then, a combination of potential global exposure image frames corresponding to adjacent shutter exposure image frames is generated based on the initial global exposure image frames. That is, in the image frame prediction part of this embodiment, adjacent shutter exposure image frames ( and and global deformation field DF As input, the output is a series of potential global exposure image frames. .
[0029] like Figure 3 The diagram shown is a schematic diagram of the principle of obtaining an initial global exposure image frame provided in this embodiment. In an optional implementation of this embodiment, the step of converting adjacent roller shutter exposure image frames into initial global exposure image frames based on the global deformation field includes: calculating the warp weight corresponding to the adjacent roller shutter exposure image frame based on a preset weight calculation formula; performing a warp operation on the adjacent roller shutter exposure image frame and the corresponding warp weight to obtain two initial global exposure image frames corresponding to the adjacent roller shutter exposure image frames.
[0030] The above weight calculation formula can be expressed as: ; in, Represents the global deformation field. This represents the first mask corresponding to the first shutter exposure image frame in an adjacent shutter exposure image frame. This represents the second mask corresponding to the second shutter exposure image frame in an adjacent shutter exposure image frame. This represents the warp weight corresponding to the first shutter exposure image frame. This represents the warp weight corresponding to the second shutter exposure image frame. This indicates a multiplication or addition operation.
[0031] It should be noted that in this embodiment, the shutter exposure image frame and the global exposure image frame are regarded as planes during the exposure period, and the mask describes the transformation relationship between the two planes.
[0032] Additionally, the warp operation formula can be expressed as:
[0033] in, This represents the warp operation, whose input is the image and the displacement vector of each pixel. This indicates the initial global exposure image frame corresponding to the first shutter exposure image frame. This indicates the initial global exposure image frame corresponding to the second shutter exposure image frame.
[0034] In one optional embodiment of this example, the step of generating a potential global exposure image frame combination corresponding to adjacent shutter exposure image frames based on the initial global exposure image frame includes: performing multiple fusion processes on the two initial global exposure image frames corresponding to the adjacent shutter exposure image frames based on a preset fusion model to obtain a potential global exposure image frame combination.
[0035] The above fusion model can be expressed as: ; in, , These represent the weighted residuals corresponding to the first and second exposure image frames in adjacent shutter exposure image frames, respectively. This represents the confidence level, which is a value between 0 and 1. This embodiment represents the feature extraction operation. This can be achieved using U-Net. This indicates a warp operation. This represents the warp weight corresponding to the first shutter exposure image frame. This represents the warp weight corresponding to the second shutter exposure image frame. and These represent the first and second shutter exposure image frames, respectively. This represents the global exposure image frame.
[0036] Step 104: Insert different combinations of potential global exposure image frames between corresponding adjacent shutter exposure image frames to obtain the interpolated video.
[0037] like Figure 4 The image shown is a schematic diagram of a video frame interpolation effect provided in this embodiment. Figure 4 (a) is an image frame exposed by the rolling shutter, with an exposure time of 0 to 0.5; (b) is a video frame exposed by the rolling shutter, with an exposure time of 0.5 to 1; (c) is the event data stream from time 0 to 1; (d) is the image frame of the global exposure at time 0.75 reconstructed in this embodiment; and (e) is the image frame of the rectangular frame portion after 32x interpolation from time 0.25 to time 0.75. This embodiment uses a deformation field to convert between frames exposed by the rolling shutter and the global shutter, taking into account both the influence of the rolling shutter and the nonlinear motion between frames, resulting in smoother and clearer interpolation results.
[0038] Next, an embodiment of this application also provides a refined video frame interpolation method, such as... Figure 5 The diagram shown is a detailed flowchart of the video frame interpolation method provided in this embodiment. The implementation process of the video frame interpolation method includes the following steps: Step 501: For each adjacent shutter exposure image frame in the video to be interpolated, acquire the event data stream collected in the exposure time interval of the adjacent shutter exposure image frames; Step 502: Divide the event data stream into multiple time segments, and then divide each time segment into multiple voxel frames to obtain the event matrix corresponding to the event data stream; Step 503: Traverse the event matrix, input each adjacent time segment into the optical flow estimation network to perform optical flow estimation, obtain the optical flow corresponding to each adjacent time segment, and determine the optical flow field composed of all optical flows as the global deformation field corresponding to the event data stream; Step 504: Calculate the mask of the global deformation field and the adjacent shutter exposure image frames based on the preset weight calculation formula to obtain the warp weight corresponding to the adjacent shutter exposure image frames. Step 505: Perform a warp operation on adjacent shutter exposure image frames and their corresponding warp weights to obtain two initial global exposure image frames corresponding to adjacent shutter exposure image frames. Step 506: Based on the preset fusion model, perform fusion processing on the two initial global exposure image frames corresponding to adjacent shutter exposure image frames to obtain a potential global exposure image frame combination. Step 507: Insert different potential global exposure image frame combinations into the corresponding adjacent shutter exposure image frames to obtain the interpolated video.
[0039] It should be understood that the sequence number of each step in this embodiment does not imply the order in which the steps are executed. The execution order of each step should be determined by its function and internal logic, and should not constitute a unique limitation on the implementation process of this application embodiment.
[0040] Next, one embodiment of this application also provides an event-guided video frame interpolation device, such as... Figure 6 The diagram shows a program module of a video frame interpolation device, which can be used to implement the video frame interpolation method in the aforementioned embodiments. The video frame interpolation device mainly includes: The acquisition module 601 is used to acquire the event data stream collected in the exposure time interval of each adjacent shutter exposure image frame in the video to be interpolated. Estimation module 602 is used to estimate the corresponding global deformation field based on the event data stream; The generation module 603 is used to generate a combination of potential global exposure image frames corresponding to adjacent shutter exposure image frames based on the global deformation field; wherein, the combination of potential global exposure image frames includes multiple potential global exposure image frames. The frame interpolation module 604 is used to insert different combinations of potential global exposure image frames between corresponding adjacent shutter exposure image frames to obtain the interpolated video.
[0041] In some embodiments of this example, the estimation module is specifically used to: divide the event data stream into multiple time segments, and then divide each time segment into multiple voxel frames to obtain the event matrix corresponding to the event data stream; traverse the event matrix, input each adjacent time segment into the optical flow estimation network to perform optical flow estimation, and obtain the optical flow corresponding to each adjacent time segment; and determine the optical flow field composed of all optical flows as the global deformation field corresponding to the event data stream.
[0042] In some embodiments of this example, the above-mentioned generation and recognition module is specifically used to: convert adjacent shutter exposure image frames into initial global exposure image frames based on the global deformation field; and generate a combination of potential global exposure image frames corresponding to adjacent shutter exposure image frames based on the initial global exposure image frames.
[0043] Furthermore, in some embodiments of this example, when the generation module performs the function of converting adjacent roller shutter exposure image frames into initial global exposure image frames based on the global deformation field, it is specifically used to: calculate the warp weights corresponding to adjacent roller shutter exposure image frames based on a preset weight calculation formula; perform warp operations on the adjacent roller shutter exposure image frames and their corresponding warp weights to obtain two initial global exposure image frames corresponding to the adjacent roller shutter exposure image frames.
[0044] Furthermore, in some other embodiments of this example, when the generation module performs the function of generating a combination of potential global exposure image frames corresponding to adjacent shutter exposure image frames based on the initial global exposure image frames, it is specifically used to: perform fusion processing on the two initial global exposure image frames corresponding to adjacent shutter exposure image frames based on a preset fusion model to obtain a combination of potential global exposure image frames.
[0045] It should be noted that the video frame interpolation methods in the foregoing embodiments can all be implemented based on the video frame interpolation device provided in this embodiment. Those skilled in the art can clearly understand that, for the sake of convenience and brevity, the specific working process of the video frame interpolation device described in this embodiment can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0046] Based on the technical solution of the above embodiments of this application, for each adjacent shutter exposure image frame in the video to be interpolated, an event data stream collected within the exposure time interval of the adjacent shutter exposure image frames is acquired; the corresponding global deformation field is estimated based on the event data stream; a potential global exposure image frame combination corresponding to the adjacent shutter exposure image frames is generated according to the global deformation field; different potential global exposure image frame combinations are inserted between the corresponding adjacent shutter exposure image frames to obtain the interpolated video. Through the implementation of the solution of this application, the conversion between shutter exposure image frames and global exposure image frames is realized based on the deformation field, and the prediction of potential image frames is guided by synchronously acquired event data. The event signal can provide effective support for the estimation of inter-frame motion state, thereby improving the interpolation effect of the video acquired in the shutter exposure mode.
[0047] Figure 7 An electronic device provided in one embodiment of this application, which can be used to implement the video frame interpolation method in the foregoing embodiments, mainly includes: The system includes a memory 701, a processor 702, and a computer program 703 stored in the memory 701 and executable on the processor 702. The memory 701 and the processor 702 are connected via communication. When the processor 702 executes the computer program 703, it implements the video frame interpolation method described in the foregoing embodiments. The number of processors can be one or more.
[0048] It should be noted that memory can be internal storage units, such as hard drives or RAM; or it can be external storage devices, such as plug-in hard drives, Smart Media Cards (SMCs), Secure Digital (SD) cards, Flash Cards, etc. Furthermore, memory can include both internal storage units and external storage devices, and it can also be used to temporarily store data that has been output or will be output. It should be noted that when the processor is a neural network chip, the electronic device may not include memory; whether the electronic device needs to use memory to store the corresponding computer program depends on the type of processor.
[0049] In addition, the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), neural network chips, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0050] One embodiment of this application also provides a computer-readable storage medium, which may be disposed in the aforementioned electronic device. The computer-readable storage medium may be as described above. Figure 7 The memory in the illustrated embodiment.
[0051] The computer-readable storage medium stores a computer program that, when executed by a processor, implements the aforementioned video frame interpolation method. Furthermore, the computer-readable storage medium can also be a USB flash drive, external hard drive, read-only memory (ROM), RAM, magnetic disk, or optical disk, or any other medium capable of storing program code.
[0052] It should be noted that the apparatuses and methods disclosed in the several embodiments provided in this application can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0053] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0054] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0055] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned readable storage medium includes various media capable of storing program code, such as USB flash drives, external hard drives, ROM, RAM, magnetic disks, or optical disks.
[0056] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0057] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0058] The above is a description of the event-guided video frame interpolation method, apparatus, device, and storage medium provided in this application. For those skilled in the art, based on the ideas of the embodiments of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An event-guided video frame interpolation method, characterized in that, include: For each adjacent shutter exposure image frame in the video to be interpolated, acquire the event data stream collected within the exposure time interval of the adjacent shutter exposure image frames; Estimate the corresponding global deformation field based on the event data stream; A potential global exposure image frame combination corresponding to the adjacent roller shutter exposure image frames is generated based on the global deformation field; wherein, the potential global exposure image frame combination includes multiple potential global exposure image frames; Different combinations of potential global exposure image frames are inserted between the corresponding adjacent shutter exposure image frames to obtain a video with complete frame interpolation.
2. The video frame interpolation method according to claim 1, characterized in that, The step of estimating the corresponding global deformation field based on the event data stream includes: The event data stream is divided into multiple time segments, and each time segment is then divided into multiple voxel frames to obtain the event matrix corresponding to the event data stream. The event matrix is traversed, and each adjacent time segment is input into the optical flow estimation network to estimate the optical flow, thereby obtaining the optical flow corresponding to each adjacent time segment. The optical flow field composed of all the optical flows is determined as the corresponding global deformation field of the event data stream.
3. The video frame interpolation method according to claim 1, characterized in that, The step of generating a combination of potential global exposure image frames corresponding to the adjacent shutter exposure image frames based on the global deformation field includes: Based on the global deformation field, the adjacent roller shutter exposure image frames are converted into initial global exposure image frames; Based on the initial global exposure image frame, a potential global exposure image frame combination corresponding to the adjacent shutter exposure image frame is generated.
4. The video frame interpolation method according to claim 3, characterized in that, The step of converting the adjacent shutter exposure image frames into an initialized global exposure image frame based on the global deformation field includes: The warp weights corresponding to the adjacent shutter exposure image frames are calculated based on a preset weight calculation formula; Perform a warp operation on the adjacent shutter exposure image frames and the corresponding warp weights to obtain two initial global exposure image frames corresponding to the adjacent shutter exposure image frames.
5. The video frame interpolation method according to claim 3, characterized in that, The step of generating a combination of potential global exposure image frames corresponding to the adjacent shutter exposure image frames based on the initialized global exposure image frames includes: Based on a preset fusion model, the two initial global exposure image frames corresponding to the adjacent roller shutter exposure image frames are fused multiple times to obtain a potential global exposure image frame combination.
6. The video frame interpolation method according to claim 5, characterized in that, The fusion model is expressed as follows: ; in, , These represent the weighted residuals corresponding to the first and second roller shutter exposure image frames in the adjacent roller shutter exposure image frames, respectively. Indicates the confidence level. This indicates a feature extraction operation. This indicates a warp operation. This represents the warp weight corresponding to the first shutter exposure image frame. This represents the warp weight corresponding to the second shutter exposure image frame. and These represent the first shutter exposure image frame and the second shutter exposure image frame, respectively. Indicates the global exposure image frame; This indicates the initial global exposure image frame corresponding to the first shutter exposure image frame. This indicates the initial global exposure image frame corresponding to the second shutter exposure image frame.
7. An event-guided video frame interpolation device, characterized in that, include: The acquisition module is used to acquire event data streams collected within the exposure time interval of each adjacent shutter exposure image frame in the video to be interpolated. The estimation module is used to estimate the corresponding global deformation field based on the event data stream; A generation module is configured to generate a combination of potential global exposure image frames corresponding to the adjacent shutter exposure image frames based on the global deformation field; wherein, the combination of potential global exposure image frames includes multiple potential global exposure image frames; The frame interpolation module is used to insert different combinations of potential global exposure image frames between corresponding adjacent roller shutter exposure image frames to obtain a frame-interpolated video.
8. An electronic device, characterized in that, Includes memory and processor, of which: The processor is used to execute computer programs stored in the memory; When the processor executes the computer program, it implements the steps in the video frame interpolation method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the video frame interpolation method according to any one of claims 1 to 6.