Video generation method, apparatus, system, electronic device, and readable storage medium
By acquiring and fusing image frames with different exposure methods, utilizing the high spatial information of short-exposure image frames and the temporal information of long-exposure image frames, and combining neural network processing, the problem of unclear video quality generated from single-frame encoded images was solved, and higher-definition video generation was achieved.
Patent Information
- Application Number
- CN202210712026.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2042-06-22
AI Technical Summary
The existing technology generates videos with insufficient clarity from single-frame encoded images, failing to meet users' visual needs, mainly because the underdetermined inverse problem has not been effectively solved.
The system acquires a first short-exposure image frame, a long-exposure image frame, and a second short-exposure image frame captured sequentially by an image capture device. The long-exposure image frame is reconstructed and fused with the two short-exposure image frames. The system generates a reconstructed frame using the high spatial information of the short-exposure image frame and the temporal information of the long-exposure image frame. A neural network is then used for refinement to improve video clarity.
The reconstructed frames generated through fusion processing have higher clarity, improving video quality and providing users with a better visual experience.
Smart Images

Figure CN115118974B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of image processing, and more particularly, to a video generation method, device, system, electronic device and readable storage medium. BACKGROUND
[0002] Video compression sensing technology can generate a video from a single frame coded image. This technology can enable users to achieve the effect of high-speed video shooting using a low-speed camera.
[0003] However, the quality of the video generated by the single frame coded image is usually not satisfactory to the user, for example, the images in the generated video are usually blurred and cannot clearly display details. This is because the process of generating a video from a single frame coded image needs to solve an underdetermined inverse problem. In the underdetermined inverse problem, the number of known conditions is usually much smaller than the number of unknowns, so it is difficult to effectively solve the unknowns. In order to solve this problem, the commonly used methods include adding artificially prescribed regularization terms, such as total variation, or introducing statistical prior information in the data set through a neural network to impose constraints on the results, so as to improve the quality of the generated video.
[0004] However, since the above methods do not actually change the underdetermined degree of the problem, they are limited by the problem, and the images in the video generated based on the single frame coded image are still not clear enough, making it difficult for users to obtain a good visual experience.
[0005] Therefore, there is a need for a new video generation method to solve the above problems. SUMMARY
[0006] To solve the above problems, the present disclosure provides a video generation method, device, system, electronic device and readable storage medium, which can improve the clarity of the images in the video generated based on the single frame coded image, and further improve the quality of the video, so that users can obtain a good visual experience.
[0007] According to an aspect of the present disclosure, a video generation method is provided, including: obtaining a first short-exposure image frame, a long-exposure image frame, and a second short-exposure image frame captured in sequence by an image capturing device, wherein the first short-exposure image frame and the second short-exposure image frame are short-exposure image frames obtained in a first coded exposure mode, and the long-exposure image frame is a single-frame coded long-exposure image frame obtained by superimposing a plurality of image frames obtained by continuous exposure in a second coded exposure mode different from the first coded exposure mode; reconstructing the long-exposure image frame to obtain a plurality of pre-reconstruction frames; for each pre-reconstruction frame in the plurality of pre-reconstruction frames, fusing the first short-exposure image frame, the second short-exposure image frame, and the each pre-reconstruction frame to generate a reconstruction frame; and generating a reconstruction video based on the plurality of reconstruction frames corresponding to the plurality of pre-reconstruction frames.
[0008] According to some embodiments of the present disclosure, wherein the first coded exposure mode is to code captured scene information with a spatially uniform modulation pattern, and the second coded exposure mode is to code scene information captured at N consecutive time instants with N spatially non-uniform modulation patterns different from each other to obtain N image frames, and superimpose the N image frames to generate the long-exposure image frame, wherein N is an integer greater than or equal to 2.
[0009] According to some embodiments of the present disclosure, wherein for each pre-reconstruction frame in the plurality of pre-reconstruction frames, fusing the first short-exposure image frame, the second short-exposure image frame, and the each pre-reconstruction frame to generate a reconstruction frame includes: for each pre-reconstruction frame in the plurality of pre-reconstruction frames, determining a first interpolation image frame between the each pre-reconstruction frame and the first short-exposure image frame, and a second interpolation image frame between the each pre-reconstruction frame and the second short-exposure image frame; fusing the first interpolation image frame, the second interpolation image frame, and the each pre-reconstruction frame to generate a reconstruction frame.
[0010] According to some embodiments of the present disclosure, wherein the determining, for each of the plurality of pre-reconstructed frames, a first interpolated image frame between the each pre-reconstructed frame and the first short-exposure image frame and a second interpolated image frame between the each pre-reconstructed frame and the second short-exposure image frame comprises: determining, for each of the plurality of pre-reconstructed frames, a first set of relative positional relationship information of objects in the each pre-reconstructed frame and corresponding objects in the first short-exposure image frame and a second set of relative positional relationship information of objects in the each pre-reconstructed frame and corresponding objects in the second short-exposure image frame; performing a mapping interpolation of spatial positions on the first short-exposure image frame based on the first set of relative positional relationship information to align spatial positions of objects in the each pre-reconstructed frame and corresponding objects in the first short-exposure image frame to obtain the first interpolated image frame; performing a mapping interpolation of spatial positions on the second short-exposure image frame based on the second set of relative positional relationship information to align spatial positions of objects in the each pre-reconstructed frame and corresponding objects in the second short-exposure image frame to obtain the second interpolated image frame.
[0011] According to some embodiments of the present disclosure, wherein the first set of relative positional relationship information and / or the second set of relative positional relationship information comprises optical flow information describing a motion direction and an offset of an object.
[0012] According to some embodiments of the present disclosure, the method further comprises: inputting the first short-exposure image frame, the first set of relative position relationship information and the first interpolated image frame and the second short-exposure image frame, the second set of relative position relationship information and the second interpolated image frame into a pre-trained first neural network model for fusion to obtain a refined first set of relative position relationship information and a first set of information weights and a refined second set of relative position relationship information and a second set of information weights, wherein the first set of information weights indicates the weight of the information of the object in the first short-exposure image frame in each pre-reconstruction frame, and the second set of information weights indicates the weight of the information of the object in the second short-exposure image frame in each pre-reconstruction frame; based on the refined first set of relative position relationship information, performing spatial position mapping interpolation on the first short-exposure image frame to align the spatial positions of the corresponding objects in the each pre-reconstruction frame and the first short-exposure image frame, and multiplying the interpolated result by the corresponding first information weight in the first set of information weights to obtain a first fine interpolation image frame; based on the refined second set of relative position relationship information, performing spatial position mapping interpolation on the second short-exposure image frame to align the spatial positions of the corresponding objects in the each pre-reconstruction frame and the second short-exposure image frame, and multiplying the interpolated result by the corresponding second information weight in the second set of information weights to obtain a second fine interpolation image frame.
[0013] According to some embodiments of the present disclosure, wherein the fusion of the first interpolated image frame, the second interpolated image frame and the each pre-reconstruction frame to generate a reconstruction frame comprises: fusion of the first fine interpolation image frame and the second fine interpolation image frame and the each pre-reconstruction frame to obtain a reconstruction frame.
[0014] According to some embodiments of the present disclosure, wherein determining, for each of the plurality of pre-reconstructed frames, a first interpolated image frame between the each pre-reconstructed frame and the first short-exposure image frame and a second interpolated image frame between the each pre-reconstructed frame and the second short-exposure image frame comprises: inputting, for each of the plurality of pre-reconstructed frames, the each pre-reconstructed frame and the first short-exposure image frame and the each pre-reconstructed frame and the second short-exposure image frame into a second neural network respectively; aligning, by the second neural network, a spatial position of an object in the each pre-reconstructed frame and a corresponding object in the first short-exposure image frame to obtain the first interpolated image frame; aligning, by the second neural network, a spatial position of an object in the each pre-reconstructed frame and a corresponding object in the second short-exposure image frame to obtain the second interpolated image frame; wherein the second neural network is pre-trained and the second neural network employs deformable convolution.
[0015] According to some embodiments of the present disclosure, wherein the object comprises one of a pixel, a coding unit or a recognizable feature of an image frame.
[0016] According to some embodiments of the present disclosure, wherein fusing the first interpolated image frame, the second interpolated image frame and the each pre-reconstructed frame to generate a reconstructed frame comprises: inputting the first interpolated image frame, the second interpolated image frame and the each pre-reconstructed frame into a pre-trained third neural network to fuse to generate a reconstructed frame, wherein the third neural network is based on a neural network structure with a UNet structure.
[0017] According to some embodiments of the present disclosure, wherein a frame rate at which the image capturing device captures image frames is lower than a frame rate of the reconstructed video.
[0018] According to some embodiments of the present disclosure, wherein the image capturing device comprises optical encoding means for encoding a captured scene using different encoding exposure modes, the optical encoding means comprising a digital micromirror device (DMD) or a liquid crystal on silicon modulator (LCoS).
[0019] According to some embodiments of the present disclosure, wherein the first short-exposure image frame, the long-exposure image frame and the second short-exposure image frame are captured consecutively by the image capturing device.
[0020] According to some embodiments of the present disclosure, the method further comprises: acquiring a second long-exposure image frame and a third short-exposure image frame sequentially captured by the image capturing device after the second short-exposure image frame, wherein the third short-exposure image frame is a short-exposure image frame obtained in the first coded exposure manner, and the second long-exposure image frame is a single-frame coded long-exposure image frame obtained by superimposing a plurality of image frames obtained by continuous exposure in the second coded exposure manner; reconstructing the second long-exposure image frame to obtain a plurality of second pre-reconstruction frames; for each of the plurality of second pre-reconstruction frames, fusing the second short-exposure image frame, the third short-exposure image frame and the each second pre-reconstruction frame to generate a second reconstruction frame; generating a second reconstruction video based on a plurality of second reconstruction frames corresponding to the plurality of second pre-reconstruction frames; and combining the reconstruction video and the second reconstruction video to generate a third reconstruction video.
[0021] According to some embodiments of the present disclosure, wherein the first short-exposure image frame and the second short-exposure image frame have higher quality spatial information than the long-exposure image frame; and the long-exposure image frame has more temporal information than the first short-exposure image frame and the second short-exposure image frame.
[0022] According to another aspect of the present disclosure, a video generation apparatus is also provided, comprising: an image frame acquisition module configured to acquire a first short-exposure image frame, a long-exposure image frame and a second short-exposure image frame sequentially captured by an image capturing device, wherein the first short-exposure image frame and the second short-exposure image frame are short-exposure image frames obtained in a first coded exposure manner, and the long-exposure image frame is a single-frame coded long-exposure image frame obtained by superimposing a plurality of image frames obtained by continuous exposure in a second coded exposure manner different from the first coded exposure manner; a pre-reconstruction module configured to reconstruct the long-exposure image frame to obtain a plurality of pre-reconstruction frames; a fusion module configured to, for each of the plurality of pre-reconstruction frames, fuse the first short-exposure image frame, the second short-exposure image frame and the each pre-reconstruction frame to generate a reconstruction frame; and a reconstruction module configured to generate a reconstruction video based on a plurality of reconstruction frames corresponding to the plurality of pre-reconstruction frames.
[0023] According to some embodiments of the present disclosure, wherein the first coded exposure manner is to code captured scene information with a spatially uniform modulation pattern, and the second coded exposure manner is to code scene information captured at N consecutive time instants with N spatially non-uniform modulation patterns different from each other to obtain N image frames, and superimpose the N image frames to generate the long-exposure image frame, wherein N is an integer greater than or equal to 2.
[0024] According to some embodiments of the present disclosure, the fusion module includes an interpolation unit and a fusion reconstruction unit, wherein the interpolation unit is configured to determine, for each of the plurality of pre-reconstruction frames, a first interpolated image frame between the each pre-reconstruction frame and the first short-exposure image frame and a second interpolated image frame between the each pre-reconstruction frame and the second short-exposure image frame; and the fusion reconstruction unit is configured to fuse the first interpolated image frame, the second interpolated image frame and the each pre-reconstruction frame to generate a reconstruction frame.
[0025] According to some embodiments of the present disclosure, the interpolation unit is configured to determine, for each of the plurality of pre-reconstruction frames, a first set of relative position relationship information of objects in the each pre-reconstruction frame and corresponding objects in the first short-exposure image frame and a second set of relative position relationship information of objects in the each pre-reconstruction frame and corresponding objects in the second short-exposure image frame; perform mapping interpolation of spatial positions on the first short-exposure image frame based on the first set of relative position relationship information to align spatial positions of objects in the each pre-reconstruction frame and corresponding objects in the first short-exposure image frame to obtain the first interpolated image frame; and perform mapping interpolation of spatial positions on the second short-exposure image frame based on the second set of relative position relationship information to align spatial positions of objects in the each pre-reconstruction frame and corresponding objects in the second short-exposure image frame to obtain the second interpolated image frame.
[0026] According to some embodiments of the present disclosure, the first set of relative position relationship information and / or the second set of relative position relationship information includes optical flow information describing motion directions and offset amounts of objects.
[0027] According to some embodiments of the present disclosure, the fusion module further comprises a refinement unit configured to: input the first short-exposure image frame, the first set of relative position relationship information, and the first interpolated image frame, and the second short-exposure image frame, the second set of relative position relationship information, and the second interpolated image frame into a pre-trained first neural network model for fusion to obtain a refined first set of relative position relationship information and a first set of information weights, and a refined second set of relative position relationship information and a second set of information weights, wherein the first set of information weights indicates the weight of the information of the object in the first short-exposure image frame in each pre-reconstruction frame, and the second set of information weights indicates the weight of the information of the object in the second short-exposure image frame in each pre-reconstruction frame; based on the refined first set of relative position relationship information, perform spatial position mapping interpolation on the first short-exposure image frame to align the spatial positions of the object in each pre-reconstruction frame and the corresponding object in the first short-exposure image frame, and multiply the result after interpolation by the corresponding first information weight in the first set of information weights to obtain a first fine interpolated image frame; based on the refined second set of relative position relationship information, perform spatial position mapping interpolation on the second short-exposure image frame to align the spatial positions of the object in each pre-reconstruction frame and the corresponding object in the second short-exposure image frame, and multiply the result after interpolation by the corresponding second information weight in the second set of information weights to obtain a second fine interpolated image frame.
[0028] According to some embodiments of the present disclosure, the fusion reconstruction unit is configured to fuse the first fine interpolated image frame and the second fine interpolated image frame with each pre-reconstruction frame to obtain a reconstruction frame.
[0029] According to some embodiments of the present disclosure, the interpolation unit is configured to: for each pre-reconstruction frame in the plurality of pre-reconstruction frames, input the pre-reconstruction frame, the first short-exposure image frame, and the pre-reconstruction frame, the second short-exposure image frame, and the pre-reconstruction frame into a second neural network respectively; align the spatial positions of the object in each pre-reconstruction frame and the corresponding object in the first short-exposure image frame through the second neural network to obtain the first interpolated image frame; align the spatial positions of the object in each pre-reconstruction frame and the corresponding object in the second short-exposure image frame through the second neural network to obtain the second interpolated image frame; wherein the second neural network is pre-trained and the second neural network adopts deformation convolution.
[0030] According to some embodiments of the present disclosure, the object comprises one of a pixel, a coding unit, or an identifiable feature of an image frame.
[0031] According to some embodiments of the present disclosure, the fusion reconstruction unit is configured to input the first interpolated image frame, the second interpolated image frame, and each of the pre-reconstructed frames into a third pre-trained neural network for fusion to generate a reconstructed frame, wherein the third neural network is based on a neural network structure having a UNet structure.
[0032] According to some embodiments of the present disclosure, a frame rate at which the image capturing device captures image frames is lower than a frame rate of the reconstructed video.
[0033] According to some embodiments of the present disclosure, the image capturing device comprises optical encoding devices for encoding a captured scene using different coded exposure modes, the optical encoding devices comprising a digital micromirror device (DMD) or a liquid crystal on silicon modulator (LCoS).
[0034] According to some embodiments of the present disclosure, the first short-exposure image frame, the long-exposure image frame, and the second short-exposure image frame are captured consecutively by the image capturing device.
[0035] According to some embodiments of the present disclosure, the image frame acquisition module is further configured to acquire a second long-exposure image frame and a third short-exposure image frame captured by the image capturing device in sequence after the second short-exposure image frame, wherein the third short-exposure image frame is a short-exposure image frame obtained in the first coded exposure mode, and the second long-exposure image frame is a single-frame coded long-exposure image frame obtained by superimposing a plurality of image frames obtained by continuous exposure in the second coded exposure mode; the pre-reconstruction module is further configured to reconstruct the second long-exposure image frame to obtain a plurality of second pre-reconstructed frames; the fusion module is further configured to, for each of the plurality of second pre-reconstructed frames, fuse the second short-exposure image frame, the third short-exposure image frame, and the each of the second pre-reconstructed frames to generate a second reconstructed frame; the reconstruction module is further configured to generate a second reconstructed video based on the plurality of second reconstructed frames corresponding to the plurality of second pre-reconstructed frames; and the video generation apparatus further comprises a combination module configured to combine the reconstructed video and the second reconstructed video to generate a third reconstructed video.
[0036] According to some embodiments of the present disclosure, the first short-exposure image frame and the second short-exposure image frame have higher quality spatial information than the long-exposure image frame; and the long-exposure image frame has more temporal information than the first short-exposure image frame and the second short-exposure image frame.
[0037] According to another aspect of the present disclosure, a video generation system is also provided, comprising: an optical encoding device configured to set a plurality of encoding exposure modes for a scene to be photographed in response to a driving signal; an image capture sensor configured to sequentially expose to capture a first short-exposure image frame, a long-exposure image frame, and a second short-exposure image frame in response to the driving signal, wherein the first short-exposure image frame and the second short-exposure image frame are short-exposure image frames obtained in a first encoding exposure mode, and the long-exposure image frame is a single-frame encoding long-exposure image frame obtained by superimposing a plurality of image frames obtained by continuous exposure in a second encoding exposure mode different from the first encoding exposure mode; and an image processor configured to reconstruct the long-exposure image frame to obtain a plurality of pre-reconstruction frames, fuse the first short-exposure image frame, the second short-exposure image frame, and each of the plurality of pre-reconstruction frames to generate a reconstruction frame for each of the plurality of pre-reconstruction frames, and generate a reconstruction video based on the plurality of reconstruction frames corresponding to the plurality of pre-reconstruction frames.
[0038] According to another aspect of the present disclosure, an electronic device is also provided, comprising: a processor; and a memory, wherein the memory has stored therein computer-readable code that, when executed by the processor, implements the video generation method described above.
[0039] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium is also provided, having stored therein computer-readable instructions that, when executed by a processor, implement the video generation method described above.
[0040] The embodiments of the present disclosure provide a video generation method, device, system, electronic device, and readable storage medium.
[0041] Therefore, according to the method of the embodiments of the present disclosure, in addition to obtaining the single-frame encoding long-exposure image frame captured in the second encoding exposure mode, two short-exposure image frames captured in the first encoding exposure mode different from the second encoding exposure mode before and after the long-exposure image frame are also obtained, by utilizing the higher spatial information in the short-exposure image frames and the more temporal information in the long-exposure image frame, the two short-exposure image frames are fused with each of the plurality of pre-reconstruction frames generated from the long-exposure image frame to generate a plurality of reconstruction frames with higher definition after optimization, and the quality of the video generated based on such reconstruction frames can be improved, thereby enabling users to obtain a good visual experience. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings needed to be used in the description of the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some of the example embodiments of the present disclosure, and other drawings can also be obtained by those of ordinary skill in the art without creative labor on the basis of these drawings.
[0043] Figure 1A is a flow chart illustrating a video generation method according to a first embodiment of the present disclosure;
[0044] Figure 1B is a schematic diagram illustrating a video generation method according to the first embodiment of the present disclosure;
[0045] Figure 2 is an example illustrating a part of an image capturing device according to an embodiment of the present disclosure;
[0046] Figure 3 is a schematic diagram illustrating the encoding of scene information to be captured using a modulation pattern to generate a long-exposure image frame;
[0047] Figure 4 is a reference Figure 3 a schematic diagram of reconstructing a long-exposure image frame to obtain a plurality of pre-reconstructed frames;
[0048] Figure 5 is a comparative effect diagram illustrating a conventional long-exposure image frame reconstruction video method and a reconstructed frame generated based on the method described in the present disclosure;
[0049] Figure 6 is a flow chart illustrating a video generation method according to a second embodiment of the present disclosure;
[0050] Figure 7 is a schematic diagram illustrating a video generation method according to the second embodiment of the present disclosure;
[0051] Figure 8 is a schematic diagram illustrating a video generation method according to a third embodiment of the present disclosure;
[0052] Figure 9 is a block diagram illustrating a video generation apparatus according to a fourth embodiment of the present disclosure;
[0053] Figure 10 is a block diagram illustrating a video generation apparatus according to a fifth embodiment of the present disclosure;
[0054] Figure 11 is a block diagram illustrating a video generation apparatus according to a sixth embodiment of the present disclosure;
[0055] Figure 12is a block diagram illustrating a video generation system according to a seventh embodiment of the present disclosure;
[0056] Figure 13 is a structural diagram illustrating an electronic device according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0057] In order to make the objects, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure. Based on the described embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative effort belong to the scope of protection of the present disclosure.
[0058] Unless otherwise defined, technical terms or scientific terms used in the present disclosure should be understood as having the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs. The terms “first”, “second”, and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are used to distinguish different components. The terms “include”, “contain”, and similar terms mean that the elements or objects before the terms encompass the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms “connect” or “connected” and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms “up”, “down”, “left”, “right”, and the like are used only to indicate relative positional relationships, and when the absolute positions of the described objects are changed, the relative positional relationships can also be changed accordingly. In order to keep the following description of the embodiments of the present disclosure clear and concise, the present disclosure omits the detailed description of some known functions and known components.
[0059] Flowcharts are used in the present disclosure to illustrate the steps of the methods according to the embodiments of the present disclosure. It should be understood that the preceding or subsequent steps do not necessarily proceed in sequence. On the contrary, various steps can be processed in reverse order or simultaneously. Meanwhile, other operations can be added to these processes, or a step or several steps can be removed from these processes.
[0060] In the specification and drawings of the present disclosure, elements are described in singular or plural form according to the embodiments. However, the selection of singular and plural forms for the proposed case is only for the convenience of explanation and is not intended to limit the present disclosure to this. Therefore, the singular form can include the plural form, and the plural form can also include the singular form, unless the context clearly indicates otherwise.
[0061] <First Embodiment>
[0062] Figure 1A A flowchart of a video generation method according to a first embodiment of the present disclosure is shown. Figure 1B A schematic diagram of a video generation method according to a first embodiment of the present disclosure is shown. In addition to acquiring a single-frame coded long-exposure image frame captured in a second coded exposure mode, the method also acquires two short-exposure image frames captured before and after the long-exposure image frame in a first coded exposure mode different from the second coded exposure mode. By utilizing the higher spatial information in the short-exposure image frames and the more temporal information in the long-exposure image frame, the two short-exposure image frames are fused with each of the plurality of pre-reconstruction frames generated from the long-exposure image frame to generate a plurality of optimized reconstruction frames with higher definition. The quality of the video generated based on such reconstruction frames can be improved, breaking through the constraints of the quality of the video based on the single-frame coded long-exposure image frame, so that the user obtains a good visual experience. The video generation method according to the present disclosure will be described in detail below with reference to Figure 1A and Figure 1B The video generation method according to the present disclosure includes the following steps:
[0063] At step S110, a first short-exposure image frame, a long-exposure image frame, and a second short-exposure image frame sequentially captured by an image capture device are acquired, wherein the first short-exposure image frame and the second short-exposure image frame are short-exposure image frames obtained in a first coded exposure mode, and the long-exposure image frame is a single-frame coded long-exposure image frame obtained by superimposing a plurality of image frames obtained by continuous exposure in a second coded exposure mode different from the first coded exposure mode. For example, in Figure 1B , a relatively clear first short-exposure image frame encoded using the first coded exposure mode is obtained at time 1. That is, the first short-exposure image frame has more spatial information. On the other hand, since the exposure time of the short-exposure image frame is short, the temporal information it contains is relatively less. A plurality of image frames used to generate the long-exposure image frame are obtained at times 2 to N-1, which are obtained by continuous exposure in a second coded exposure mode different from the first coded exposure mode, wherein the second coded exposure mode uses different modulation patterns for each time. These image frames are superimposed together to generate a long-exposure image frame in a long-exposure process. However, the spatial information of these image frames is limited by the reconstruction quality when they are encoded. Then a relatively clear second short-exposure image frame encoded using the first coded exposure mode is obtained at time N. The characteristics of the second short-exposure image frame can refer to the first short-exposure image frame. From Figure 1B It can be seen that the objects in the figure are moving over time from time 1 to time N, and these image frames have relatively more temporal information due to the longer exposure time.
[0064] At step S120, the long-exposure image frame is reconstructed to obtain a plurality of pre-reconstruction frames.
[0065] At step S130, for each of the plurality of pre-reconstructed frames, the first short-exposure image frame, the second short-exposure image frame and the each pre-reconstructed frame are fused to generate a reconstructed frame. As shown, the fusion utilizes the higher spatial information in the short-exposure image frames and the more temporal information in the long-exposure image frames, such that the generated plurality of reconstructed frames are clearer. Figure 1B
[0066] At step S140, a reconstructed video is generated based on the plurality of reconstructed frames corresponding to the plurality of pre-reconstructed frames.
[0067] In particular, first, at step S110, a first short-exposure image frame, a long-exposure image frame and a second short-exposure image frame captured sequentially by an image capturing device can be obtained, wherein the first short-exposure image frame and the second short-exposure image frame can be short-exposure image frames obtained in a first coded exposure manner, and the long-exposure image frame can be a single-frame coded long-exposure image frame obtained by superimposing a plurality of image frames obtained by continuous exposure in a second coded exposure manner different from the first coded exposure manner.
[0068] In one example, the image capturing device can be used to capture image frames. The image capturing device can be any type of device capable of capturing images, such as a camera, a video camera, a smartphone, a tablet computer, a laptop computer, or a fixed or portable device with image capturing function, etc. In this document, image frames, images, frames, etc. can be used interchangeably.
[0069] In the video compression sensing based technology, the superimposition of a plurality of coded frames to generate a single-frame coded long-exposure image frame can be achieved by encoding the scene information to be captured in a specific coded exposure manner when capturing images using the image capturing device. Generally, the coded exposure manner refers to encoding the scene information to be captured using a modulation pattern or mask set by an optical coding device. According to examples of the present disclosure, the image capturing device can include an optical coding device for encoding the captured scene using different coded exposure manners. The optical coding device can include a digital micromirror device (DMD), a liquid crystal on silicon modulator (LCoS), or other optical devices capable of setting a modulation pattern or mask. In one example, the optical coding device can also be referred to as a spatial light modulator. In the present disclosure, both modulation patterns and masks are used to encode the scene information to be captured, so modulation patterns and masks can be used interchangeably herein.
[0070] Figure 2 is an example showing a part of an image capturing device according to an embodiment of the present disclosure. As shown, the image capturing device includes an optical coding device 1000 and an image sensor 2000. The optical coding device 1000 is configured to set a modulation pattern or mask to encode the scene information to be captured. The image sensor 2000 is configured to capture the encoded scene information to generate image frames. In one example, the image sensor 2000 can be a CMOS sensor or a CCD sensor. Figure 2 As shown, one example of the image capture device can include an objective lens, a DMD, a relay lens, and an image sensor, where the image sensor can include an image sensor such as a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS). In addition, the DMD can also be replaced by other optical encoding devices. The objective lens and the relay lens are conventional components for transmitting optical information, which are not described here again. In one example, the image capture device can control the optical encoding device it contains to encode the scene to be captured, and control the corresponding image sensor to expose and image to generate an encoded image frame. For example, a driving signal can be used to control the optical encoding device and the image sensor. The driving signal can be generated by the image capture device or by other external devices coupled to the image capture device.
[0071] In one example, the image capture device can obtain a short-exposure image frame in a first encoding exposure manner at a single time instance, and can obtain a plurality of image frames in a second encoding exposure manner and superimpose them to obtain a single-frame encoded long-exposure image frame. The first and the second are only to represent that the short-exposure image frame and the long-exposure image frame have different encoding exposure manners. For a plurality of short-exposure image frames and / or long-exposure image frames, the encoding exposure manner of each short-exposure image frame and / or long-exposure image frame can be different from each other. For example, for two long-exposure image frames, their encoding exposure manners can be different.
[0072] In the long-exposure image frame, the continuous exposure to obtain a plurality of image frames means that the scene information is encoded at a plurality of time instances within one exposure to obtain a plurality of image frames. For example, when encoding a long-exposure image frame, within one exposure of the image sensor, the optical encoding device can dynamically refresh the modulation pattern multiple times in response to the driving signal, and each modulation pattern corresponds to the scene information at the current time instance; at the same time, the image sensor can capture the scene information encoded by each modulation pattern according to the driving signal, and then the plurality of encoded image frames are superimposed along the time dimension within one long-exposure process to obtain a single-frame encoded long-exposure image frame. In one example, the encoding exposure manner for the long-exposure image frame is different for each time instance of the scene, so the encoding manner can also be referred to as time-varying encoding with spatial structure.
[0073] Figure 3 A schematic diagram is shown for encoding the scene information to be captured using a modulation pattern to generate a long-exposure image frame. As shown, the image capture device can include an objective lens, a DMD, a relay lens, and an image sensor. The image sensor can include an image sensor such as a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS). In addition, the DMD can also be replaced by other optical encoding devices. The objective lens and the relay lens are conventional components for transmitting optical information, which are not described here again. In one example, the image capture device can control the optical encoding device it contains to encode the scene to be captured, and control the corresponding image sensor to expose and image to generate an encoded image frame. For example, a driving signal can be used to control the optical encoding device and the image sensor. The driving signal can be generated by the image capture device or by other external devices coupled to the image capture device. Figure 3As shown, according to embodiments of the present disclosure, the encoding exposure manner for generating a long-exposure image frame is to encode scene information captured at consecutive N time instants with N mutually different spatially non-uniform modulation patterns to obtain N image frames, and to superimpose the N image frames to generate one long-exposure image frame, where N is an integer greater than or equal to 2. It is appreciated by those skilled in the art that based on the video compressive sensing technology, a long-exposure image frame can be generated by the encoding exposure manner, and such a long-exposure image frame can be reconstructed into a video with multiple image frames based on the encoding exposure manner, and thus the specific content of encoding and decoding image frames with modulation patterns will not be elaborated here.
[0074] In one example, since the image capturing device has optical encoding modulation devices, the captured short-exposure image frames will also be encoded by the modulation patterns in general. According to embodiments of the present disclosure, for the short-exposure image frames, the captured scene information can be encoded with a spatially uniform modulation pattern. Such a modulation pattern is usually a global constant pattern, so that the scene spatial information in the short-exposure image frames can be completely preserved.
[0075] In one example, the exposure duration of the short-exposure image frames is less than the total exposure duration of the long-exposure image frames. For example, the short-exposure image frames can have a minimum unit exposure duration that can be supported by the image capturing device, while the long-exposure image frames have 8 or 16 minimum unit exposure durations. However, it is appreciated by those skilled in the art that the exposure duration can be set according to actual needs, and is not limited to a fixed exposure duration. For example, the short-exposure image frames can have 2 minimum unit exposure durations. In addition, multiple short-exposure image frames and / or long-exposure image frames can also each have different exposure durations. For example, a first long-exposure image frame can have 8 minimum unit exposure durations, while a second long-exposure image frame can have 16 minimum unit exposure durations.
[0076] In one example, the method according to the present disclosure requires acquiring at least three exposure image frames, i.e. a first short exposure image frame, a long exposure image frame and a second short exposure image frame. Among them, the first short exposure image frame, the long exposure image frame and the second short exposure image frame are captured in time sequence. For example, the first short exposure image frame is captured at time 1, the long exposure image frame is captured at time 2 to N-1, and the second short exposure image frame is captured at time N. It should be noted that the above image frames acquired can be discontinuous, for example, the second short exposure image frame is not captured at time N but at time N+1. Since there is a time domain correlation between the image frames for video, there is usually more information available for reference between two closely adjacent image frames, so it will be more effective to use continuously captured images. Preferably, according to the embodiment of the present disclosure, the first short exposure image frame, the long exposure image frame and the second short exposure image frame are continuously captured by the image capture device.
[0077] In one example, after the image frames are captured by the image capture device, since the captured image frames can also have, for example, a mark indicating the capture time, the order of acquiring the image frames can be acquired in any order without being based on the capture order.
[0078] After acquiring the first short exposure image frame, the long exposure image frame and the second short exposure image frame, at step S120, the long exposure image frame can be reconstructed to obtain a plurality of pre-reconstruction frames.
[0079] Based on the video compressive sensing technology, the long exposure image frame can be reconstructed to obtain a plurality of pre-reconstruction frames for subsequent processing by using the encoding exposure mode adopted when capturing the long exposure image frame. Here, a neural network (for example, a convolutional neural network based on residual block) can be generally used to implement the reconstruction of the long exposure image frame, and the reconstruction algorithm used can also be based on, for example, GAP-TV, E2E-CNN, DUN and any other applicable algorithm.
[0080] Figure 4 is a reference Figure 3 The schematic diagram of reconstructing the long exposure image frame to obtain a plurality of pre-reconstruction frames. As Figure 4 shown, using Figure 3 the N spatially non-uniform modulation patterns adopted in the method of the present disclosure to process the video can obtain N pre-reconstruction frames.
[0081] After obtaining a plurality of pre-reconstruction frames, at step S130, for each of the plurality of pre-reconstruction frames, the first short exposure image frame, the second short exposure image frame and the each pre-reconstruction frame can be fused to generate a reconstruction frame.
[0082] According to embodiments of the present disclosure, the first interpolated image frame, the second interpolated image frame and each pre-reconstruction frame can be input into a pre-trained neural network for fusion to generate a reconstruction frame, where the neural network can be based on a neural network structure with a UNet structure, such as a conventional U-Net, a RA-UNet, a Swin-Conv-UNet, or the neural network can be based on other applicable neural network structures. A neural network with a UNet structure generally has cross-layer connections and spatial down-sampling, and such a network structure can more effectively fuse image frames.
[0083] According to embodiments of the present disclosure, a short-exposure image frame generally has higher quality spatial information compared to a long-exposure image frame, while a long-exposure image frame has more temporal information compared to a short-exposure image frame. The fusion in the method described in the present disclosure optimizes a pre-reconstruction frame using the higher quality spatial information of two short-exposure image frames near the pre-reconstruction frame, so that the generated reconstruction frame has higher clarity.
[0084] Figure 5 is a comparative effect diagram showing the reconstruction video method based on a conventional long-exposure image frame and the reconstruction frame generated based on the method described in the present disclosure. In Figure 5 , the first row of images shows the first short-exposure image frame, the long-exposure image frame and the second long-exposure image frame captured in sequence by the image device. The second row of images shows part of the reconstruction frames (reconstruction frames 3, 9, 15) in the video generated based on the prior art. Among them, the fourth image in the second row shows the magnified effect of the area surrounded by the black frame in the reconstruction frame 15 based on the prior art. In addition, the third row of images shows part of the reconstruction frames in the video generated based on the method of the present disclosure. Among them, the fourth image in the third row shows the magnified effect of the area surrounded by the black frame in the reconstruction frame 15 based on the method of the present disclosure. It can be seen that the clarity of the reconstruction frame generated based on the method described in the present disclosure has a significant improvement compared to the reconstruction frame generated based on the prior art.
[0085] Finally, at step S140, a reconstruction video can be generated based on the plurality of reconstruction frames corresponding to the plurality of pre-reconstruction frames. According to embodiments of the present disclosure, the frame rate at which the image capturing device captures image frames can be lower than the frame rate of the reconstruction video, thereby enabling a user to achieve the effect of high-speed video shooting using a low-speed camera.
[0086] In addition, on the basis of the video generation method based on the first embodiment, a sequence of short-exposure image frames and long-exposure image frames can be captured alternately in sequence, so that a continuous longer reconstructed video can be generated, and a better visual experience can be provided for the user. According to an embodiment of the present disclosure, a second long-exposure image frame and a third short-exposure image frame captured in sequence by the image capturing device after the second short-exposure image frame can be obtained, wherein the third short-exposure image frame is a short-exposure image frame obtained in the first coded exposure manner, and the second long-exposure image frame is a single-frame coded long-exposure image frame obtained by superimposing a plurality of image frames obtained by continuous exposure in the second coded exposure manner; the second long-exposure image frame is reconstructed to obtain a plurality of second pre-reconstruction frames; for each second pre-reconstruction frame in the plurality of second pre-reconstruction frames, the second short-exposure image frame, the third short-exposure image frame, and each second pre-reconstruction frame are fused to generate a second reconstruction frame; a second reconstruction video is generated based on the plurality of second reconstruction frames corresponding to the plurality of second pre-reconstruction frames; and the reconstructed video and the second reconstruction video are combined to generate a third reconstruction video.
[0087] In addition, in one example, the image capturing device such as the captured image frames or other data can be stored in a storage device, which can include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing data.
[0088] In one example, the data in the image capturing device can also be transmitted to a server via a wired network or a wireless network for other devices to obtain or directly transmitted to other devices, and then further processed by other devices on the captured images. In one example, the network can be a wired network and / or a wireless network. For example, the wired network can transmit data in the form of twisted pair, coaxial cable or optical fiber transmission, and the wireless network can transmit data in the form of 3G / 4G / 5G mobile communication network, Bluetooth, Zigbee or WiFi.
[0089] In one example, the image processor for processing image frames can also be directly or indirectly coupled to the image capturing device to obtain the image frames captured by the image capturing device. In one example, the image processor can be a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a ready-to-program field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc. which can perform image processing functions. In one example, the image processor can also be integrated with the image capturing device, such as a camera comprising the image capturing device and the image processor.
[0090] The above is combined with Fig. 1- Figure 5 The video generation method described in the present disclosure is described in detail. As can be known from the above detailed description, according to the first embodiment of the present disclosure, in addition to acquiring a single-frame coded long-exposure image frame captured in a second coded exposure mode, two short-exposure image frames captured in a first coded exposure mode different from the second coded exposure mode before and after the long-exposure image frame are also acquired, by utilizing the higher spatial information in the short-exposure image frames and the more temporal information in the long-exposure image frame, the two short-exposure image frames are fused with each of the plurality of pre-reconstruction frames generated from the long-exposure image frame to generate a plurality of optimized reconstruction frames with higher clarity, and the quality of the video generated based on such reconstruction frames can be improved, thereby providing users with a good visual experience.
[0091] <Second embodiment>
[0092] Figure 6 is a flowchart showing a video generation method according to the second embodiment of the present disclosure. In Figure 6 In the method shown, the steps in the first embodiment are further optimized to achieve reconstruction frames with higher clarity, thereby improving the quality of the generated video. Figure 6 The steps shown in are similar to those shown in Fig. 1, so the same reference numerals are used to mark them and will not be described here.
[0093] As Figure 6 As shown in the second embodiment of the present disclosure, after generating a plurality of pre-reconstruction frames based on steps S110 and S120, step S130 can include step S610: for each of the plurality of pre-reconstruction frames, determining a first interpolation image frame between each of the pre-reconstruction frames and the first short-exposure image frame and a second interpolation image frame between each of the pre-reconstruction frames and the second short-exposure image frame; and step S620: fusing the first interpolation image frame, the second interpolation image frame, and each of the pre-reconstruction frames to generate a reconstruction frame.
[0094] The interpolated image frame can be obtained by spatial position mapping interpolation based on the motion information and the offset of the object in the image, can be obtained by processing the image frame using a neural network with deformation convolution, or can be obtained in other manners based on the information of the object in the two image frames. In one example, the image processor can perform corresponding operations in units of objects in the image. According to one embodiment of the present disclosure, the object in the image includes at least one of the following: a pixel of the image frame, a coding unit, a recognizable feature, or a processable unit capable of representing the motion information and the offset in the image frame in other manners. The recognizable feature can be a specific feature in the image recognized based on image recognition technology. For example, the recognizable feature can be a plane as shown in FIG. 1B. Figure 1B
[0095] In the scheme of spatial position mapping interpolation, according to one embodiment of the present disclosure, for each of the plurality of pre-reconstruction frames, a first set of relative position relationship information of the object in each pre-reconstruction frame and the corresponding object in the first short-exposure image frame and a second set of relative position relationship information of the object in each pre-reconstruction frame and the corresponding object in the second short-exposure image frame can be determined; based on the first set of relative position relationship information, the first short-exposure image frame is spatially positionally interpolated to align the spatial positions of the object in each pre-reconstruction frame and the corresponding object in the first short-exposure image frame, to obtain a first interpolated image frame; based on the second set of relative position relationship information, the second short-exposure image frame is spatially positionally interpolated to align the spatial positions of the object in each pre-reconstruction frame and the corresponding object in the second short-exposure image frame, to obtain a second interpolated image frame.
[0096] According to one embodiment of the present disclosure, the first set of relative position relationship information and / or the second set of relative position relationship information includes optical flow information describing the motion direction and the offset of the object. When optical flow information is used herein, the optical flow information can be considered as optical flow information for a single object in the image frame or a set of optical flow information for multiple or all objects in the image according to the context. In one example, the method of calculating the optical flow information can be implemented based on existing technologies such as PWCNet, RAFT, AMP, etc., and thus will not be described herein. In one example, the set of relative position relationship information can also be other information describing the motion direction and the offset of the object obtained according to existing technologies.
[0097] In one example, the first interpolated image frame, the second interpolated image frame, and each pre-reconstruction frame can be input into a pre-trained neural network for fusion to generate the reconstruction frame in a similar manner as the first embodiment, where the neural network is based on a neural network structure with a UNet structure, such as a conventional U-Net, a RA-UNet, a Swin-Conv-UNet, or the neural network can be based on other applicable neural network structures.
[0098] Taking the optical flow information as an example, Figure 7 is a schematic diagram illustrating a video generation method according to a second embodiment of the present disclosure. As Figure 7 shown, at step S710, steps like S110 are performed to obtain a first short-exposure image frame, a long-exposure image frame, and a second short-exposure image frame. At step S720, steps like S120 are performed to obtain N pre-reconstruction frames.
[0099] Further, at step S730, for the kth pre-reconstruction frame among the N pre-reconstruction frames, first optical flow information of the kth pre-reconstruction frame to the first short-exposure image frame and second optical flow information of the kth pre-reconstruction frame to the second short-exposure image frame are calculated; then the first short-exposure image frame is mapped and interpolated in spatial position based on the first optical flow information to align the spatial positions of the objects in the kth pre-reconstruction frame and the corresponding objects in the first short-exposure image frame to obtain a first interpolated image frame for the kth pre-reconstruction frame, and the second short-exposure image frame is mapped and interpolated in spatial position based on the second optical flow information to align the spatial positions of the objects in the kth pre-reconstruction frame and the corresponding objects in the second short-exposure image frame to obtain a second interpolated image frame for the kth pre-reconstruction frame; then the first interpolated image frame, the second interpolated image frame, and the kth pre-reconstruction frame are fused based on a neural network to generate a reconstruction frame corresponding to the kth pre-reconstruction frame.
[0100] Finally, at step 740, a reconstruction image is generated based on the N reconstruction frames for the N pre-reconstruction frames.
[0101] The above is an interpolation method implemented based on one-way optical flow information of the reconstruction frame to the short-exposure image frame. According to some examples of the present disclosure, the reconstruction frame can also be generated based on two-way optical flow information, i.e., both the optical flow information of the reconstruction frame to the short-exposure image frame and the optical flow information of the short-exposure image frame to the reconstruction frame. The method based on two-way optical flow information is similar to the method described with respect to Figure 7 The two-way optical flow information can make the spatial positions of the objects in the interpolated image frame closer to the positions of the corresponding objects in the pre-reconstruction frame, thus the clarity of the generated reconstruction frame can be further improved compared to using one-way optical flow information. In addition, compared to one-way optical flow information, two-way optical flow information will generally consume more computing resources.
[0102] According to the second embodiment of the present disclosure, the interpolated image frame can also be obtained without the set of spatial relative position relationship information between image frames. According to the embodiments of the present disclosure, for each of the plurality of pre-reconstruction frames, each pre-reconstruction frame and the first short-exposure image frame and each pre-reconstruction frame and the second short-exposure image frame are respectively input into a neural network; the spatial positions of the objects in the first pre-reconstruction frame and the corresponding objects in the first short-exposure image frame are aligned by the neural network to obtain a first interpolated image; the spatial positions of the objects in the first pre-reconstruction frame and the second short-exposure image frame are aligned by the neural network to obtain a second interpolated image; wherein the neural network is pre-trained and the neural network adopts deformation convolution.
[0103] The neural network adopting deformation convolution can directly realize the spatial displacement processing of the objects in the image frames to generate the interpolated image frames by using the information in the image frames, without actually calculating the relative position relationship, motion direction and offset of the objects in the image frames.
[0104] According to the second embodiment of the present disclosure, compared with the spatial positions of the objects in the short-exposure, the spatial positions of the objects in the interpolated frame are more consistent with the spatial positions of the corresponding objects in the pre-reconstruction frame, so that the fusion of the interpolated frame and the pre-reconstruction frame further improves the definition of the reconstructed frame, thereby further improving the quality of the image.
[0105] <Third Embodiment>
[0106] On the basis of the second embodiment of the present disclosure, the method of the third embodiment of the present disclosure can further include further optimization of the interpolated image frame to further improve the definition of the reconstructed frame.
[0107] According to an embodiment of the present disclosure, after obtaining the first interpolation image frame and the second interpolation image frame for each pre-reconstruction frame, the first short-exposure image frame, the first relative position relationship information set and the first interpolation image frame and the second short-exposure image frame, the second relative position relationship information set and the second interpolation image frame can also be input into a pre-trained neural network model for fusion to obtain a refined first relative position relationship information set and a first information weight set and a refined second relative position relationship information set and a second information weight set, wherein the first information weight set indicates the weight of the information of the object in the first short-exposure image frame in each pre-reconstruction frame, and the second information weight set indicates the weight of the information of the object in the second short-exposure image frame in each pre-reconstruction frame; based on the refined first relative position relationship information set, the first short-exposure image frame can be mapped and interpolated in spatial position to align the spatial positions of the object in each pre-reconstruction frame and the corresponding object in the first short-exposure image frame, and the result after interpolation is multiplied by the corresponding first information weight in the first information weight set to obtain a first fine interpolation image; based on the refined second relative position relationship information set, the second short-exposure image frame can be mapped and interpolated in spatial position to align the spatial positions of the object in each pre-reconstruction frame and the corresponding object in the second short-exposure image frame, and the result after interpolation is multiplied by the corresponding second information weight in the second information weight set to obtain a second fine interpolation image.
[0108] According to an embodiment of the present disclosure, after obtaining the first fine interpolation image frame and the second fine interpolation image frame, the first fine interpolation image frame and the second fine interpolation image frame and each pre-reconstruction frame can be fused to obtain a reconstruction frame.
[0109] Specifically, Figure 8 is a schematic diagram illustrating a video generation method according to a third embodiment of the present disclosure. In Figure 8 In step S810, steps like S110, S710 are performed to obtain the first short-exposure image frame, the long-exposure image frame and the second short-exposure image frame. In step S820, steps like S120, S720 are performed to obtain N pre-reconstruction frames.
[0110] Further, in step S830, first, the first interpolation image frame and the second interpolation image frame are obtained based on a method similar to step S730.
[0111] It should be noted that the method of obtaining the first interpolation image frame and the second interpolation image frame can be based on the relative position relationship information set or based on a neural network with deformation convolution. However, since the relative position relationship information set is usually needed when further optimizing the interpolation image frame, it is preferred to obtain the first interpolation image frame and the second interpolation image frame based on the relative position relationship information set.
[0112] In one example, the set of relative position relationship information is taken as the first optical flow information and the second optical flow information shown in the above formula (1) and (2), but the set of relative position relationship information is not limited to the first optical flow information and the second optical flow information. Figure 7
[0113] Then, for the kth pre-reconstruction frame, the first short-exposure image frame, the first optical flow information and the first interpolated image frame, and the second short-exposure image frame, the second optical flow information and the second interpolated image frame are input into the pre-trained neural network model for fusion to obtain the refined first optical flow information and the first set of information weights and the refined second relative optical flow information and the second set of information weights, wherein the first set of information weights indicates the weight of the information of the object in the first short-exposure image frame in the kth pre-reconstruction frame, and the second set of information weights indicates the weight of the information of the object in the second short-exposure image frame in the kth pre-reconstruction frame, wherein the set of weight information can include a weight map; based on the refined first optical flow information, the first short-exposure image frame is mapped and interpolated in spatial position to align the spatial positions of the object in the kth pre-reconstruction frame and the corresponding object in the first short-exposure image frame, and the result after interpolation is multiplied by the corresponding first information weight in the first set of information weights to obtain a first fine interpolation image; based on the refined second optical flow information, the second short-exposure image frame is mapped and interpolated in spatial position to align the spatial positions of the object in the kth pre-reconstruction frame and the corresponding object in the second short-exposure image frame, and the result after interpolation is multiplied by the corresponding second information weight in the second set of information weights to obtain a second fine interpolation image; then the first fine interpolation image, the second fine interpolation image and the kth pre-reconstruction frame are fused based on the neural network to generate a reconstruction frame corresponding to the kth pre-reconstruction frame.
[0114] Finally, in step S840, a reconstruction image is generated based on the N reconstruction frames for the N pre-reconstruction frames in a manner similar to steps S140 and S740.
[0115] According to the third embodiment of the present disclosure, compared with the interpolation image frame, the object in the fine interpolation image frame is further refined, so that the spatial positions of the object in the fine interpolation image frame and the object in the pre-reconstruction frame are more consistent, and therefore the clarity of the image frame obtained by fusing the fine interpolation frame and the pre-reconstruction frame is further improved.
[0116] <Fourth Embodiment>
[0117] In addition to providing the above-mentioned video generation method, the present disclosure also provides a video generation device 900, which will be described in detail below in combination with the following drawings. Figure 9 This will be described in detail.
[0118] Figure 9 is a block diagram illustrating a video generation apparatus 900 according to a fourth embodiment of the present disclosure. As shown in Figure 9 the video generation apparatus 900 described in the present disclosure can include an image frame acquisition module 910, a pre-reconstruction module 920, a fusion module 930, and a reconstruction module 940.
[0119] According to an embodiment of the present disclosure, the image frame acquisition module 910 can be configured to acquire a first short-exposure image frame, a long-exposure image frame, and a second short-exposure image frame captured by an image capturing device in sequence, wherein the first short-exposure image frame and the second short-exposure image frame are short-exposure image frames obtained in a first coded exposure manner, and the long-exposure image frame is a single-frame coded long-exposure image frame obtained by superimposing a plurality of image frames obtained by continuous exposure in a second coded exposure manner different from the first coded exposure manner.
[0120] According to an embodiment of the present disclosure, the first coded exposure manner for generating the short-exposure image frame can be to code the captured scene information with a spatially uniform modulation pattern.
[0121] According to an embodiment of the present disclosure, the second coded exposure manner for generating the long-exposure image frame can be to code the scene information captured at N consecutive time instants with N spatially non-uniform modulation patterns different from each other to obtain N image frames, and to superimpose the N image frames to generate the long-exposure image frame, wherein N is an integer greater than or equal to 2.
[0122] According to an embodiment of the present disclosure, the short-exposure image frame generally has higher quality spatial information than the long-exposure image frame, and the long-exposure image frame has more temporal information than the short-exposure image frame.
[0123] According to an embodiment of the present disclosure, the first short-exposure image frame, the long-exposure image frame, and the second short-exposure image frame can be captured by the image capturing device in sequence.
[0124] According to an example of the present disclosure, the image capturing device can include an optical coding device for coding the captured scene with different coded exposure manners. The optical coding device can include a digital micromirror device (DMD), a liquid crystal on silicon modulator (LCoS), or other optical devices capable of setting a modulation pattern or mask.
[0125] According to an embodiment of the present disclosure, the pre-reconstruction module 920 can be configured to reconstruct the long-exposure image frame to obtain a plurality of pre-reconstruction frames.
[0126] According to an embodiment of the present disclosure, the fusion module 930 can be configured to, for each of the plurality of pre-reconstruction frames, fuse the first short-exposure image frame, the second short-exposure image frame, and the each pre-reconstruction frame to generate a reconstruction frame.
[0127] According to an embodiment of the present disclosure, the first interpolated image frame, the second interpolated image frame, and each of the pre-reconstructed frames can be input into a pre-trained neural network for fusion to generate the reconstructed frame, where the neural network can be based on a neural network structure with a UNet structure, such as a conventional U-Net, a RA-UNet, a Swin-Conv-UNet, or the neural network can be based on other applicable neural network structures.
[0128] According to an embodiment of the present disclosure, the reconstruction module 940 can be configured to generate a reconstructed video based on a plurality of reconstructed frames corresponding to the plurality of pre-reconstructed frames.
[0129] According to an embodiment of the present disclosure, the frame rate at which the image capturing device captures the image frames can be lower than the frame rate of the reconstructed video.
[0130] According to an embodiment of the present disclosure, the image frame acquisition module 910 can be further configured to acquire a second long-exposure image frame and a third short-exposure image frame captured by the image capturing device in sequence after the second short-exposure image frame, where the third short-exposure image frame is a short-exposure image frame obtained in the first coded exposure manner, and the second long-exposure image frame is a single-frame coded long-exposure image frame obtained by superimposing a plurality of image frames obtained by continuous exposure in the second coded exposure manner; the pre-reconstruction module 920 can be further configured to reconstruct the second long-exposure image frame to obtain a plurality of second pre-reconstructed frames; the fusion module 930 can be further configured to, for each of the plurality of second pre-reconstructed frames, fuse the second short-exposure image frame, the third short-exposure image frame, and each of the second pre-reconstructed frames to generate a second reconstructed frame; the reconstruction module 940 can be further configured to generate a second reconstructed video based on a plurality of second reconstructed frames corresponding to the plurality of second pre-reconstructed frames; and the video generation apparatus 900 can further include a combination module configured to combine the reconstructed video and the second reconstructed video to generate a third reconstructed video.
[0131] <5th Embodiment>
[0132] Figure 10 is a block diagram illustrating a video generation apparatus 1000 according to a fifth embodiment of the present disclosure, where Figure 10 has some components in common with Figure 9 some components shown in
[0133] As shown in Figure 10 , the fusion module 930 can include an interpolation unit 1010 and a fusion reconstruction unit 1020.
[0134] According to an embodiment of the present disclosure, the interpolation unit 1010 can be configured to determine, for each of the plurality of pre-reconstruction frames, a first interpolation image frame between the each pre-reconstruction frame and the first short-exposure image frame and a second interpolation image frame between the each pre-reconstruction frame and the second short-exposure image frame; and the fusion reconstruction unit 1020 can be configured to fuse the first interpolation image frame, the second interpolation image frame and the each pre-reconstruction frame to generate a reconstruction frame.
[0135] According to an embodiment of the present disclosure, the interpolation unit 1010 can be further configured to determine, for each of the plurality of pre-reconstruction frames, a first set of relative position relationship information of objects in the each pre-reconstruction frame and corresponding objects in the first short-exposure image frame and a second set of relative position relationship information of the objects in the each pre-reconstruction frame and corresponding objects in the second short-exposure image frame; perform mapping interpolation of spatial positions on the first short-exposure image frame based on the first set of relative position relationship information to align spatial positions of the objects in the first pre-reconstruction frame and the corresponding objects in the first short-exposure image frame to obtain the first interpolation image; and perform mapping interpolation of spatial positions on the second short-exposure image frame based on the second set of relative position relationship information to align spatial positions of the objects in the first pre-reconstruction frame and the corresponding objects in the second short-exposure image frame to obtain the second interpolation image.
[0136] According to an embodiment of the present disclosure, the first set of relative position relationship information and / or the second set of relative position relationship information can include optical flow information describing a motion direction and an offset of the objects.
[0137] According to an embodiment of the present disclosure, the interpolation unit 1010 can be further configured to input, for each of the plurality of pre-reconstruction frames, the each pre-reconstruction frame and the first short-exposure image frame and the each pre-reconstruction frame and the second short-exposure image frame into a neural network respectively; align spatial positions of the objects in the first pre-reconstruction frame and the corresponding objects in the first short-exposure image frame through the neural network to obtain the first interpolation image; and align spatial positions of the objects in the first pre-reconstruction frame and the corresponding objects in the second short-exposure image frame through the neural network to obtain the second interpolation image; wherein the neural network is pre-trained and the neural network adopts deformation convolution.
[0138] According to an embodiment of the present disclosure, the objects can include one of pixels, coding units or identifiable features of the image frames.
[0139] <Sixth Embodiment>
[0140] Figure 11 is a block diagram illustrating a video generation apparatus 1100 according to a sixth embodiment of the present disclosure, wherein Figure 11 parts of Figure 9 ,Figure 10 The parts shown are identical to those shown in FIG. 1, and are therefore shown with the same reference numerals and will not be described again.
[0141] As shown in FIG. 9, the fusion module 930 can further include a refinement unit 1110. Figure 11
[0142] According to an embodiment of the present disclosure, the refinement unit 1110 can be configured to input the first short-exposure image frame, the first set of relative position relationship information, and the first interpolated image frame into a pre-trained neural network model for fusion to obtain a refined first set of relative position relationship information and a first set of information weights, wherein the first set of information weights indicates the weight of the information of the object in the first short-exposure image frame in each pre-reconstruction frame; input the second short-exposure image frame, the second set of relative position relationship information, and the second interpolated image frame into the neural network model for fusion to obtain a refined second set of relative position relationship information and a second set of information weights, wherein the second set of information weights indicates the weight of the information of the object in the second short-exposure image frame in each pre-reconstruction frame; perform mapping interpolation of the spatial position of the first short-exposure image frame based on the refined first set of relative position relationship information to align the spatial positions of the object in each pre-reconstruction frame and the corresponding object in the first short-exposure image frame, and multiply the result after interpolation by the corresponding first information weight in the first set of information weights to obtain a first fine interpolation image; perform mapping interpolation of the spatial position of the second short-exposure image frame based on the refined second set of relative position relationship information to align the spatial positions of the object in each pre-reconstruction frame and the corresponding object in the second short-exposure image frame, and multiply the result after interpolation by the corresponding second information weight in the second set of information weights to obtain a second fine interpolation image.
[0143] According to an embodiment of the present disclosure, the fusion reconstruction unit 1020 can be configured to fuse the first fine interpolation image frame and the second fine interpolation image frame and each pre-reconstruction frame to obtain a reconstruction frame.
[0144] <Seventh Embodiment>
[0145] In addition to providing the video generation method and device described above, the present disclosure also provides a video generation system. Next, the video generation system according to the seventh embodiment of the present disclosure will be described in detail. Figure 12 with reference to the accompanying drawings.
[0146] Figure 12 FIG. 12 is a block diagram illustrating a video generation system 1200 according to the seventh embodiment of the present disclosure. As shown in FIG. 12, the video generation system 1200 according to the present disclosure can include an optical encoding device 1210, an image capture sensor 1220, and an image processor 1230. Figure 9
[0147] According to an embodiment of the present disclosure, the optical encoding device 1210 can be configured to set a plurality of encoding exposure modes for a scene to be photographed in response to a driving signal.
[0148] According to an embodiment of the present disclosure, the image capturing sensor 1220 can be configured to sequentially expose to capture a first short-exposure image frame, a long-exposure image frame, and a second short-exposure image frame in response to the driving signal, wherein the first short-exposure image frame and the second short-exposure image frame can be short-exposure image frames obtained in a first encoding exposure mode, and the long-exposure image frame can be a single-frame encoding long-exposure image frame obtained by superimposing a plurality of image frames obtained by continuous exposure in a second encoding exposure mode different from the first encoding exposure mode.
[0149] According to an embodiment of the present disclosure, the image processor 1230 can be configured to reconstruct the long-exposure image frame to obtain a plurality of pre-reconstruction frames; for each of the plurality of pre-reconstruction frames, fuse the first short-exposure image frame, the second short-exposure image frame, and each of the plurality of pre-reconstruction frames to generate a reconstruction frame; and generate a reconstruction video based on the plurality of reconstruction frames corresponding to the plurality of pre-reconstruction frames.
[0150] In one example, the video generation system 1200 can be any device having a capturing, encoding exposure, and image processing function, such as a camera. In one example, the video generation system 1200 can further include an optical device, such as a lens, to capture scene information. In one example, the video generation system 1200 can further include a driving circuit that can generate a driving signal to drive the optical encoding device 1210 and the image capturing sensor 1220. In one example, the video generation system 1200 can further include an input / output (I / O) component. The I / O component can also be directly or indirectly coupled to the video generation system 1200. The I / O component can represent or interact with a modem, a keyboard, a mouse, a touch screen, or the like. In some cases, the I / O component can be implemented as part of the processor. In some cases, a user can interact with the system 1200 via the I / O component or via a hardware component controlled by the I / O component. In one example, the video generated by the image processor can be imaged to the user via the I / O component. In one example, the user can adjust the encoding exposure mode of the image capturing sensor 1220 or adjust the parameters of the image capturing sensor 1220, etc. via the I / O component.
[0151] Regarding Figures 9 to 12 Some specific details of the video generation apparatus and system shown can also refer to the content of the video generation method shown in FIGS. 1 to Figure 8 .
[0152] Figure 13is a structural diagram illustrating an electronic device 1300 according to some embodiments of the present disclosure.
[0153] Referring to Figure 13 The electronic device 1300 can include a processor 1301 and a memory 1302. The processor 1301 and the memory 1302 can be connected through a bus 1303. The electronic device 1300 can be any type of portable device (such as a smart camera, a smartphone, a tablet, etc.) or any type of fixed device (such as a desktop computer, a server, etc.).
[0154] The processor 1301 can perform various actions and processes according to programs stored in the memory 1302. Specifically, the processor 1301 can be an integrated circuit chip having a processing capability of signals. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a ready programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can be any conventional processor, etc., which can be X86 architecture or ARM architecture.
[0155] The memory 1302 stores computer executable instructions that, when executed by the processor 801, implement the above-mentioned video generation method. The memory 1302 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM can be used, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DRAM). It should be noted that the memory of the method described herein is intended to include, but not limited to, these and any other suitable types of memory.
[0156] Further, a video generation method according to the present disclosure can be recorded in a computer-readable recording medium. Specifically, according to the present disclosure, a computer-readable recording medium storing computer executable instructions, which, when executed by a processor, can cause the processor to perform a video generation method as described above, can be provided.
[0157] It should be noted that the flowchart and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or combinations of hardware and software.
[0158] In general, the various example embodiments of the present disclosure can be implemented in hardware or special-purpose circuits, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software which can be executed by a controller, microprocessor or other computing device, Although the various aspects of embodiments of the present disclosure can be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein can be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controler or other computing devices, or some combination thereof.
[0159] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0160] The foregoing is a summary of the disclosure, and shall not be considered a limitation as to the scope thereof. While several exemplary embodiments of the disclosure have been described, it will be apparent to those of ordinary skill in the art that many modifications are possible without departing from the teachings of the disclosure, which are intended to be limited only by the terms of the claims. It is to be understood that the above description is intended to be illustrative, and not restrictive. Many other embodiments will be apparent to those of skill in the art upon reviewing the above description. The scope of the disclosure should, therefore, be determined not with reference to the above description, but instead with reference to the appended claims, along with their full scope of equivalents.
Claims
1. A method for generating a video, the method comprising: obtaining a first short-exposure image frame, a long-exposure image frame and a second short-exposure image frame captured sequentially by an image capturing device, wherein the first short-exposure image frame and the second short-exposure image frame are short-exposure image frames obtained in a first coded exposure manner, and the long-exposure image frame is a single-frame coded long-exposure image frame obtained by superimposing a plurality of image frames obtained in a second coded exposure manner different from the first coded exposure manner; reconstructing the long-exposure image frame to obtain a plurality of pre-reconstructed frames; for each of the plurality of pre-reconstructed frames, determining a first set of relative positional relationship information of objects in the each pre-reconstructed frame and corresponding objects in the first short-exposure image frame, and a second set of relative positional relationship information of objects in the each pre-reconstructed frame and corresponding objects in the second short-exposure image frame; based on the first set of relative positional relationship information, performing a spatial position mapping interpolation on the first short-exposure image frame to align the spatial positions of objects in the each pre-reconstructed frame and corresponding objects in the first short-exposure image frame, to obtain a first interpolated image frame between the each pre-reconstructed frame and the first short-exposure image frame; based on the second set of relative positional relationship information, performing a spatial position mapping interpolation on the second short-exposure image frame to align the spatial positions of objects in the each pre-reconstructed frame and corresponding objects in the second short-exposure image frame, to obtain a second interpolated image frame between the each pre-reconstructed frame and the second short-exposure image frame; fusing the first interpolated image frame, the second interpolated image frame and the each pre-reconstructed frame to generate a reconstructed frame; and generating a reconstructed video based on a plurality of reconstructed frames corresponding to the plurality of pre-reconstructed frames. the first coded exposure manner is to code the captured scene information with a spatially uniform modulation pattern, and 2. The video generation method of claim 1, wherein, the second coded exposure manner is to code the scene information captured at N different time instants with N spatially non-uniform modulation patterns to obtain N image frames, and superimpose the N image frames to generate the long-exposure image frame, wherein N is an integer greater than or equal to 2. the first set of relative positional relationship information and / or the second set of relative positional relationship information comprises optical flow information describing the motion direction and offset of objects.
3. The video generation method of claim 1, wherein, 4. The method of claim 1, further comprising: inputting the first short-exposure image frame, the first set of relative position relationship information, and the first interpolation image frame and the second short-exposure image frame, the second set of relative position relationship information, and the second interpolation image frame into a first neural network model that is pre-trained and fused to obtain a refined first set of relative position relationship information and a first set of information weights and a refined second set of relative position relationship information and a second set of information weights, wherein the first set of information weights indicates a weight of information of an object in the first short-exposure image frame in each pre-reconstruction frame, and the second set of information weights indicates a weight of information of an object in the second short-exposure image frame in each pre-reconstruction frame; based on the refined first set of relative position relationship information, performing spatial position mapping interpolation on the first short-exposure image frame to align spatial positions of an object in each pre-reconstruction frame and a corresponding object in the first short-exposure image frame, and multiplying a result after interpolation by a corresponding first information weight in the first set of information weights to obtain a first fine interpolation image frame; based on the refined second set of relative position relationship information, performing spatial position mapping interpolation on the second short-exposure image frame to align spatial positions of an object in each pre-reconstruction frame and a corresponding object in the second short-exposure image frame, and multiplying a result after interpolation by a corresponding second information weight in the second set of information weights to obtain a second fine interpolation image frame.
5. The video generation method of claim 4, wherein, fusing the first interpolation image frame, the second interpolation image frame, and each pre-reconstruction frame to generate a reconstruction frame includes: fusing the first fine interpolation image frame and the second fine interpolation image frame and each pre-reconstruction frame to obtain a reconstruction frame.
6. The video generation method of claim 1, wherein, For each pre-reconstruction frame in the plurality of pre-reconstruction frames, determining a first interpolation image frame between the pre-reconstruction frame and the first short-exposure image frame and a second interpolation image frame between the pre-reconstruction frame and the second short-exposure image frame includes: For each pre-reconstruction frame in the plurality of pre-reconstruction frames, inputting the pre-reconstruction frame and the first short-exposure image frame and the pre-reconstruction frame and the second short-exposure image frame into a second neural network, respectively; aligning, by the second neural network, spatial positions of an object in each pre-reconstruction frame and a corresponding object in the first short-exposure image frame to obtain the first interpolation image frame; aligning, by the second neural network, spatial positions of an object in each pre-reconstruction frame and a corresponding object in the second short-exposure image frame to obtain the second interpolation image frame; wherein the second neural network is pre-trained and the second neural network adopts deformation convolution.
7. The video generation method of any one of claims 1-6, wherein, The object includes one of a pixel, a coding unit, or an identifiable feature of an image frame.
8. The video generation method of any one of claims 1-6, wherein, fusing the first interpolation image frame, the second interpolation image frame, and each pre-reconstruction frame to generate a reconstruction frame includes: inputting the first interpolated image frame, the second interpolated image frame and each of the pre-reconstructed frames into a third pre-trained neural network for fusion to generate a reconstructed frame, wherein the third neural network is based on a neural network structure having a UNet structure.
9. The video generation method of any one of claims 1-6, wherein, The frame rate at which the image capturing device captures image frames is lower than the frame rate of the reconstructed video.
10. The video generation method of any one of claims 1-6, wherein, The image capturing device comprises optical encoding means for encoding a captured scene using different coded exposure modes, the optical encoding means comprising a digital micromirror device (DMD) or a liquid crystal on silicon modulator (LCoS).
11. The video generation method of any one of claims 1-3, wherein, The first short-exposure image frame, the long-exposure image frame and the second short-exposure image frame are captured consecutively by the image capturing device.
12. The video generation method of any one of claims 1-3, further comprising: obtaining a second long-exposure image frame and a third short-exposure image frame captured by the image capturing device consecutively after the second short-exposure image frame, wherein the third short-exposure image frame is a short-exposure image frame obtained with the first coded exposure mode and the second long-exposure image frame is a single-frame coded long-exposure image frame obtained by superimposing a plurality of image frames obtained with the second coded exposure mode consecutively exposed; reconstructing the second long-exposure image frame to obtain a plurality of second pre-reconstructed frames; for each of the plurality of second pre-reconstructed frames, fusing the second short-exposure image frame, the third short-exposure image frame and the each of the second pre-reconstructed frames to generate a second reconstructed frame; generating a second reconstructed video based on the plurality of second reconstructed frames corresponding to the plurality of second pre-reconstructed frames; and combining the reconstructed video and the second reconstructed video to generate a third reconstructed video. The first short-exposure image frame and the second short-exposure image frame have higher quality spatial information than the long-exposure image frame; and 13. The video generation method of any one of claims 1-6, wherein, The long-exposure image frame has more temporal information than the first short-exposure image frame and the second short-exposure image frame.
14. A video generation apparatus, comprising: an image frame obtaining module configured to obtain a first short-exposure image frame, a long-exposure image frame and a second short-exposure image frame captured consecutively by an image capturing device, wherein the first short-exposure image frame and the second short-exposure image frame are short-exposure image frames obtained with a first coded exposure mode and the long-exposure image frame is a single-frame coded long-exposure image frame obtained by superimposing a plurality of image frames obtained with a second coded exposure mode different from the first coded exposure mode consecutively exposed; a pre-reconstruction module configured to reconstruct the long-exposure image frame to obtain a plurality of pre-reconstructed frames; a fusion module, wherein the fusion module comprises an interpolation unit and a fusion reconstruction unit, wherein The interpolation unit is configured to determine, for each of the plurality of pre-reconstruction frames, a first set of relative positional relationship information of objects in the each pre-reconstruction frame and corresponding objects in the first short-exposure image frame and a second set of relative positional relationship information of objects in the each pre-reconstruction frame and corresponding objects in the second short-exposure image frame; perform mapping interpolation of spatial positions on the first short-exposure image frame based on the first set of relative positional relationship information to align spatial positions of objects in the each pre-reconstruction frame and corresponding objects in the first short-exposure image frame, to obtain a first interpolated image frame between the each pre-reconstruction frame and the first short-exposure image frame; perform mapping interpolation of spatial positions on the second short-exposure image frame based on the second set of relative positional relationship information to align spatial positions of objects in the each pre-reconstruction frame and corresponding objects in the second short-exposure image frame, to obtain a second interpolated image frame between the each pre-reconstruction frame and the second short-exposure image frame; The fusion reconstruction unit is configured to fuse the first interpolated image frame, the second interpolated image frame and the each pre-reconstruction frame to generate a reconstruction frame; The reconstruction module is configured to generate a reconstruction video based on a plurality of reconstruction frames corresponding to the plurality of pre-reconstruction frames.
15. The video generation device of claim 14, wherein, The first encoding exposure mode is to encode captured scene information with a spatially uniform modulation pattern, and The second encoding exposure mode is to encode scene information captured at consecutive N time instants with N spatially non-uniform modulation patterns to obtain N image frames, and to superimpose the N image frames to generate the long-exposure image frame, where N is an integer greater than or equal to 2.
16. The video generation apparatus of claim 14, wherein the first set of relative positional relationship information and / or the second set of relative positional relationship information comprises optical flow information describing a motion direction and an offset of an object.
17. The video generation apparatus of claim 14, wherein, The fusion module further comprises a refinement unit configured to: input the first short-exposure image frame, the first set of relative positional relationship information and the first interpolated image frame and the second short-exposure image frame, the second set of relative positional relationship information and the second interpolated image frame into a pre-trained first neural network model for fusion, to obtain a refined first set of relative positional relationship information and a first set of information weights and a refined second set of relative positional relationship information and a second set of information weights, wherein the first set of information weights indicates a weight of information of an object in the first short-exposure image frame in the each pre-reconstruction frame, and the second set of information weights indicates a weight of information of an object in the second short-exposure image frame in the each pre-reconstruction frame; map and interpolate the spatial positions of the first short-exposure image frame based on the refined first set of relative positional relationship information, so as to align the spatial positions of the objects in each pre-reconstruction frame and the corresponding objects in the first short-exposure image frame, and multiply the result after the interpolation by the corresponding first information weight in the first set of information weights to obtain a first fine-interpolation image frame; map and interpolate the spatial positions of the second short-exposure image frame based on the refined second set of relative positional relationship information, so as to align the spatial positions of the objects in each pre-reconstruction frame and the corresponding objects in the second short-exposure image frame, and multiply the result after the interpolation by the corresponding second information weight in the second set of information weights to obtain a second fine-interpolation image frame.
18. The video generation device of claim 17, wherein, The fusion reconstruction unit is configured to: fuse the first fine-interpolation image frame, the second fine-interpolation image frame, and each pre-reconstruction frame to obtain a reconstruction frame.
19. The video generation device of claim 14, wherein, The interpolation unit is configured to: input each pre-reconstruction frame and the first short-exposure image frame and each pre-reconstruction frame and the second short-exposure image frame into a second neural network respectively for each pre-reconstruction frame in the plurality of pre-reconstruction frames; align the spatial positions of the objects in each pre-reconstruction frame and the corresponding objects in the first short-exposure image frame through the second neural network to obtain the first interpolation image frame; align the spatial positions of the objects in each pre-reconstruction frame and the corresponding objects in the second short-exposure image frame through the second neural network to obtain the second interpolation image frame; The second neural network is pre-trained and adopts deformation convolution.
20. The video generation apparatus of any of claims 14-19, wherein, The object includes one of a pixel, a coding unit, or an identifiable feature of an image frame.
21. The video generation apparatus of any of claims 14-19, wherein, The fusion reconstruction unit is configured to: input the first interpolation image frame, the second interpolation image frame, and each pre-reconstruction frame into a pre-trained third neural network for fusion to generate a reconstruction frame, wherein the third neural network is based on a neural network structure with a UNet structure.
22. The video generation apparatus of any of claims 14-19, wherein, The frame rate at which the image capture device captures image frames is lower than the frame rate of the reconstructed video.
23. The video generation apparatus of any of claims 14-19, wherein, The image capture device includes optical encoding devices for encoding a captured scene using different encoding exposure modes, and the optical encoding devices include a digital micromirror device (DMD) or a liquid crystal on silicon (LCoS) modulator.
24. The video generation apparatus according to any one of claims 14-19, wherein, The first short-exposure image frame, the long-exposure image frame, and the second short-exposure image frame are captured consecutively by the image capture device.
25. The video generation apparatus of any of claims 14-19, wherein, The image frame acquisition module is further configured to acquire a second long-exposure image frame and a third short-exposure image frame captured by the image capture device in sequence after the second short-exposure image frame, wherein the third short-exposure image frame is a short-exposure image frame obtained in the first encoding exposure mode, and the second long-exposure image frame is a single-frame encoding long-exposure image frame obtained by superimposing a plurality of image frames obtained by continuous exposure in the second encoding exposure mode; The pre-reconstruction module is further configured to reconstruct the second long-exposure image frame to obtain a plurality of second pre-reconstructed frames; The fusion module is further configured to, for each of the plurality of second pre-reconstructed frames, fuse the second short-exposure image frame, the third short-exposure image frame, and the each of the second pre-reconstructed frames to generate a second reconstructed frame; The reconstruction module is further configured to generate a second reconstructed video based on a plurality of second reconstructed frames corresponding to the plurality of second pre-reconstructed frames; and The video generation apparatus further comprises a combination module configured to combine the reconstructed video and the second reconstructed video to generate a third reconstructed video.
26. The video generation apparatus of any of claims 14-19, wherein, The first short-exposure image frame and the second short-exposure image frame have higher quality spatial information than the long-exposure image frame; And The long-exposure image frame has more temporal information than the first short-exposure image frame and the second short-exposure image frame.
27. A video generation system, comprising: an optical encoding device configured to set a plurality of encoding exposure modes for a scene to be photographed in response to a driving signal; an image capture sensor configured to sequentially expose to capture a first short-exposure image frame, a long-exposure image frame, and a second short-exposure image frame in response to the driving signal, wherein the first short-exposure image frame and the second short-exposure image frame are short-exposure image frames obtained in a first encoding exposure mode, and the long-exposure image frame is a single-frame encoding long-exposure image frame obtained by superimposing a plurality of image frames obtained by continuous exposure in a second encoding exposure mode different from the first encoding exposure mode; and an image processor configured to reconstruct the long-exposure image frame to obtain a plurality of pre-reconstructed frames, determine, for each of the plurality of pre-reconstructed frames, a first set of relative positional relationship information of objects in the each of the pre-reconstructed frames and corresponding objects in the first short-exposure image frame and a second set of relative positional relationship information of objects in the each of the pre-reconstructed frames and corresponding objects in the second short-exposure image frame, perform mapping interpolation of spatial positions of the first short-exposure image frame based on the first set of relative positional relationship information to align spatial positions of objects in the each of the pre-reconstructed frames and corresponding objects in the first short-exposure image frame to obtain a first interpolated image frame between the each of the pre-reconstructed frames and the first short-exposure image frame, perform mapping interpolation of spatial positions of the second short-exposure image frame based on the second set of relative positional relationship information to align spatial positions of objects in the each of the pre-reconstructed frames and corresponding objects in the second short-exposure image frame to obtain a second interpolated image frame between the each of the pre-reconstructed frames and the second short-exposure image frame, and fuse the first interpolated image frame, the second interpolated image frame, and the each of the pre-reconstructed frames to generate a reconstructed frame; and generate a reconstructed video based on a plurality of reconstructed frames corresponding to the plurality of pre-reconstructed frames.
28. An electronic device, comprising: a processor; and a memory, wherein the memory has stored thereon computer readable code which, when executed by the processor, implements the video generation method of any one of claims 1-13.
29. A non-transitory computer-readable storage medium storing computer-readable instructions, wherein, when executed by a processor, implement the video generation method of any one of claims 1-13.
Citation Information
Patent Citations
Method of filming high dynamic range videos
CN104349069A
Multi-frame interpolation method based on convolutional neural network
CN110191299A
Image processing method and device, electronic equipment and computer readable storage medium
CN111462021A