Image enhancement method, device and system
By combining frame-based and event-based sensors and utilizing event stream encoding of brightness variations for image processing, the image quality issues of motion blur and high dynamic range scenes in low-light conditions of traditional devices are solved, achieving clear image capture and simplified computation.
Patent Information
- Application Number
- CN202180044850.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-23
- Filing Date
- 2021-06-23
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2041-06-23
AI Technical Summary
Traditional image and video capture devices are prone to motion blur and artifacts in low light conditions, and image processing in high dynamic range scenes is computationally complex, which is difficult to solve effectively with existing technologies.
By combining frame-based and event-based sensors, reference image frames and event streams are acquired synchronously. Brightness changes are encoded using the event streams, and image time shifting, deblurring, reblurring, and nonlinear compensation are performed to remove artifacts and restore clear images.
It effectively removes motion blur and artifacts under low-light conditions, captures clear images of high dynamic range scenes, reduces computational complexity, and improves image quality.
Smart Images

Figure CN115836318B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to the field of image processing, in particular to image enhancement. BACKGROUND
[0002] Conventional cameras and systems for images and videos are often limited by several factors. In particular, at the sensor design level, in order to have small, low cost and better sensitive image sensors, image sensors usually use a rolling shutter readout. While this readout has advantages, it can create artifacts due to motion occurring while capturing a frame.
[0003] In addition, when a conventional camera captures a video or an image, the settings of the camera provide different balances for the acquisition, like the typical exposure / ISO balance, for example changing the exposure duration with more or less light integrated, but still with the risk of motion creating a blurred image. Keeping a small exposure duration can mitigate the motion blur, but for low light situations this will create too dark images. For the latter, it is common to compensate the small input signal by increasing the ISO gain of the sensor. However, this also amplifies the noise of the sensor.
[0004] Moreover, some scenes have both high luminosity and dark areas. For these scenes, one usually chooses between an exposure that better exposes the high luminosity and an exposure that better exposes the dark areas. A typical solution to this problem is to combine several images and use some non-linear function to compress the measured light intensity to create an HDR image. However, this solution requires the alignment of usually 2 to 8 several images and is therefore computationally too complex for high frame rate videos used for conventional image and video acquisition.
[0005] In the art, there have been some solutions that try to compensate for the above-mentioned limitations of conventional imaging and video acquisition. For example, WO2014042894A1, US10194101B1 and US9554048B2 disclose how to remove specific rolling shutter effects, WO2016107961A1, EP2454876B1 and US8687913B2 disclose how to remove blur without rolling shutter effects, and some other papers discuss the computation of high dynamic range (HDR) images and videos without removing rolling shutter and blur. However, they are usually tailored to specific artifacts and / or they have important computational complexity.
[0006] The present invention aims at improving these drawbacks. SUMMARY
[0007] In this regard, according to one aspect of the present invention, there is provided an image enhancement method comprising:
[0008] - obtaining a reference image frame of a scene by a frame-based sensor, said reference image frame containing artifacts and having an exposure duration;
[0009] - obtaining an event stream by an event-based sensor synchronized with said frame-based sensor at least during the exposure duration of said reference image frame, wherein said events encode luminance variations of the scene corresponding to said reference image; and
[0010] - deriving a corrected image frame without said artifacts from said reference image by means of said event stream.
[0011] Thus, according to this arrangement according to the invention, frame image and event data are combined to remove image artifacts, such as wobble (also called jello effect), skew, spatial aliasing, temporal aliasing, especially for scenes with motion in low light and with large dynamic range.
[0012] In an embodiment according to the invention, said reference image frame is obtained at a reference time, and the step of deriving a corrected image frame comprises a step of image time shifting, wherein said reference image frame is shifted to said corrected image frame at a target time, inside or outside said exposure duration, by establishing a relationship between said events and said reference image frame and accumulating luminance variations encoded by events obtained during a time interval [t_r, t] between said reference time and said target time, in order to shift said image at a different time, i.e. target time, when acquiring the image.
[0013] Further, according to the invention, said reference image frame is a blurred frame, and said method comprises a step of image deblurring for restoring said blurred frame to a sharp image, in order to obtain a sharp reference image, for example when said exposure duration is too large.
[0014] Alternatively, the method according to the invention comprises a step of re-blurring said corrected image frame, in order to simulate a larger exposure duration effect from a sharp image, which makes the image more realistic, especially in fast motion scenes.
[0015] Alternatively, the method according to the invention comprises a step of non-linear compensation for compensating non-linear processing both in said image frame and in said events.
[0016] Further, in another embodiment of the invention, said reference image has a first resolution and / or a first color space, and said event stream has a second resolution and / or a second color space different from said first resolution and / or said first color space, and wherein said method further comprises a step of resolution and color restoration for compensating resolution and / or color space differences between said reference image and said event stream.
[0017] In another embodiment of the application, the image frame is generated by a frame-based camera, the frame-based camera containing a rolling shutter sensor with rows of sensor elements (pixels). Further, for example, the step of deriving a corrected image frame is performed sequentially for each row of the rolling shutter sensor.
[0018] In another embodiment of the application, the method comprises generating a plurality of corrected image frames distributed at different times in the exposure duration.
[0019] Alternatively, the method according to the application comprises aligning a plurality of corrected image frames derived in different exposure durations.
[0020] According to another aspect of the application, there is further provided an image enhancement device comprising:
[0021] - a frame-based camera containing a frame-based sensor;
[0022] - an event-based camera containing an event-based sensor, wherein the frame-based camera and the event-based camera observe the same scene and generate image frames and event streams, respectively, corresponding to the scene, and the frame-based camera and the event-based camera are synchronized; and
[0023] - a processor for performing the above method.
[0024] Alternatively, according to another aspect of the application, there is further provided another image enhancement device comprising:
[0025] - a camera adapted to generate image frames and event streams corresponding to a scene;
[0026] - a processor for performing the above method.
[0027] Further, according to another aspect of the application, there is further provided an image enhancement system comprising:
[0028] - a capture module adapted to obtain a reference image which can contain artifacts and has an exposure duration, and to obtain an event stream corresponding to the scene during the exposure duration of the reference image frame;
[0029] - an image enhancement module adapted to derive from the reference image a corrected image frame free of the artifacts by means of the event stream.
[0030] In one embodiment of the system according to the application, the image enhancement module comprises an image temporal shift module for shifting the reference image frame to the corrected image frame at any time (t) inside or outside the exposure duration by accumulating the luminance variations encoded by the events.
[0031] In another embodiment of the system according to the application, the image enhancement module comprises a non-linearity compensation module for compensating non-linear processing in the image frames and in the events.
[0032] In another embodiment of the system according to the application, the reference image has a first resolution and / or a first color space and the event stream has a second resolution and / or a second color space different from the first resolution and / or the first color space, and wherein the image enhancement module further comprises a resolution and color restoration module for compensating resolution and / or color space differences between the reference image and the event stream.
[0033] The application extends the capabilities of standard frame-based image sensors by pairing them with event streams, for example from DVS sensors, in order to extend the traditional balance of image sensors and allow capturing videos and images that better match the scene characteristics. In this regard, the application allows capturing images and videos of scenes with motion in low light and with large dynamic range with limited computational complexity requirements. BRIEF DESCRIPTION OF DRAWINGS
[0034] Other features and advantages of the application will become apparent in the light of the following description and accompanying drawings in which:
[0035] Figure 1 Exposure window for images captured by frame-based sensors and for images corrected by the method according to the application is shown;
[0036] Figure 2 Functional block diagram showing a first exemplary method according to the application is shown;
[0037] Figure 3 Functional block diagram showing a second exemplary method according to the application is shown;
[0038] Figure 4 Functional block diagram showing a third exemplary method according to the application is shown;
[0039] Figure 5 Functional block diagram showing a fourth exemplary method according to the application is shown; and
[0040] Figure 6Exposure windows for two images captured by frame-based sensors and for a plurality of images corrected by the method according to the invention are shown. DETAILED DESCRIPTION
[0041] The context of the invention includes an exemplary image capture architecture comprising a conventional image frame-based camera and an event-based camera, wherein the frame-based camera and the event-based camera observe the same scene in which objects or the entire scene move with respect to the frame-based camera and the event-based camera, and the frame-based camera captures and produces image frames of the scene, including moving objects, and the event-based camera produces an event stream corresponding to the same scene including moving objects captured during (frame period) and possibly before and after the acquisition of image frames.
[0042] In particular, the frame-based camera has for example a rolling shutter sensor with rows of sensors, for example CMOS sensors, wherein the rows of sensors are rapidly scanned in sequence across the scene, row by row. Thus, when scanning from the first row to the last row of the sensor during a period called frame period, an image is captured or produced.
[0043] The event-based camera contains an asynchronous vision sensor facing the same scene observed by the frame-based camera and receives light from the same scene. The asynchronous vision sensor comprises an array of sensing elements (pixels), each sensing element including a light sensing element. Each sensing element (pixel) produces an event corresponding to a change in light in the scene. The asynchronous vision sensor can be a contrast detection (CD) sensor or a dynamic vision sensor (DVS).
[0044] Alternatively, the invention can use another possible exemplary image capture architecture comprising a single sensor recording both frames and events, for example a DAVIS sensor or an asynchronous time-based image sensor (ATIS).
[0045] In the invention, it is necessary to synchronize the frame data to the event data.
[0046] In particular, in the case of combining an image frame-based sensor with a DVS or CD sensor, it is possible to synchronize the time bases of the two sensors by exchanging synchronization signals between the sensors.
[0047] For example, one possibility is to rely on a flash signal from a frame-based camera that is usually used to trigger the flash during the exposure duration of a frame acquisition, and use this signal to convey the start of the exposure duration to the CD sensor. The CD sensor can then generate a timestamp in its own time base for the received signal, so that the processing element can relate the event timestamp to the exposure duration of the frame-based sensor.
[0048] This synchronization is important because the present invention is based on the relation between frames and events, which will be discussed below.
[0049] The events of an event-based camera (DVS or CD) are related to the temporal evolution of the luminance in the scene. An event-based sensor records "events" that encode a change in luminance with respect to a reference state. We call this reference state the reference image L(t_r) at a reference time t_r. Then, an event-based camera produces a positive event at a (pixel) position x at a time t if:
[0050] log(L(x,t) / L(x,t_r))>c+
[0051] or a negative event if:
[0052] log(L(x,t) / L(x,t_r))<c-
[0053] where c+ and c- are respectively the positive and negative threshold. Note that the use of the logarithm function here is only an example, but these relations can be adapted to any monotonic and continuous function. We also define the polarity associated with a particular event e by p e p can be written
[0054] e p = τ(log(L(x,t) / L(x,t_r)), c+, c-)
[0055] where τ(x, c+, c-) is equal to 1 if x >= c+, 0 if x ∈ (-c-, c+), and -1 if x <= -c-.
[0056] If, in a time interval T = [t_r, t_n], we receive a series of events E[t_r, t_n] = e1,..., ek at a position x, we get
[0057] ∑ e∈E[t_r,t_n] e p= τ(log(L(x, t1) / L(x, t_r)), c+, c-) +... + τ(log(L(x, tn) / L(x, t_(n-1))), c+, c-) ~ log(L(x, tn) / L(x, t_r)), c+, c-)
[0058] By defining h(x, E[t_r, t_n], c+, c-) = ∑ e+∈E[t_r,t_n],ep>0 c++∑ e-∈E[t_r,t_n],ep<0 c-, we get
[0059] (1) h(x, E[t_r, t_n], c+, c-) ~ log(L(x, tn) / L(x, t_r), c+, c-)
[0060] where equation (1) applies independently to each pixel.
[0061] In this case, equation (1) establishes a relationship between two images at different times and the events received between these times. Next, it is discussed how this relationship can be used to manipulate and improve the images according to the invention.
[0062] In the context of the invention, it also includes a processor or computing unit suitable for carrying out the image enhancement method based on the following modules.
[0063] • Image time shift module
[0064] According to equation (1), it is possible to derive the sharp image L(x, t) at an arbitrary time t from the sharp reference frame L(x, t_r) with t_r < t, by accumulating the luminance changes encoded by the events E[t_r, t] and updating the luminance of L(x, t_r), i.e.
[0065] (2) L(x, t) ~ L(x, t_r) e h(x,E[t_r,t],c+,c-)
[0066] If t_r > t, then h(x, E[t_r, t], c+, c-) = - (∑ e+∈E[t_r,t],ep>0 c++∑ e-∈E[t_r,t],ep< 0c-).
[0067] Equation (2) allows to estimate the image at an arbitrary time t, given a reference image at time t_r and the events recorded in the time interval [t_r, t]. This equation forms one of the basic modules of the method according to the invention, the image time shift module, since it allows to "shift" the images when they are acquired at different times.
[0068] • Image deblurring module
[0069] Equation (2) assumes that a reference image is available at time t_r. This assumption can be true if the exposure duration is short enough (i.e. if the image is sharp). However, if the exposure duration is too long, it is not possible to have a sharp reference image available with motion in the scene or relative motion between the scene and the camera.
[0070] Therefore, it is necessary to ensure that a sharp image at t_r is shifted. Mathematically, the blurred frame can be defined as:
[0071] B = 1 / T∫ f+T / 2 f-T / 2 L(t)dt
[0072] where B is the blurred image, f is a given time within the exposure duration and is used as a reference time, T is the exposure duration and L(t) is the sharp image at time t.
[0073] Equation (2) can be substituted at the location of L(t) and it results in:
[0074] (3) L(f) ~ T*B / ∫ f+T / 2 f-T / 2 e h(E[f,t],c+,c-) dt
[0075] Equation (3) defines an image deblurring module which allows to recover a sharp image from a blurred image at any time within the exposure duration. In the case of the image deblurring module and the image time shift module, it is possible to reconstruct a sharp image at any given time t.
[0076] • an image re-blurring module
[0077] Sometimes, a photographer or a video producer wants to keep / add a bit of blur in his images / videos. In this regard, an image can be simulated at an arbitrary exposure duration T_n by simply averaging all the possible different images that occur during the exposure. For example, if a sharp image is reconstructed at time t_r and the image is centered at time t_r with an exposure T_n, the sharp image can be computed by doing:
[0078] (4) B ~ L(t_r) / T∫ t_r+T_n / 2 t_r-T_n / 2 e h(E[t_r,t],c+,c-) dt
[0079] Equation (4) defines an image re-blurring module where a sharp image is simulated from a blurred image with a larger exposure duration.
[0080] • a resolution and color recovery module
[0081] The above defined modules act independently on each pixel. Also, the above modules assume that there is a 1 to 1 correspondence between event-based sensor pixels and frame pixels, and that c+ and c- are known for each pixel. However, this can not be the case even for combined sensors. For example, in WO 2020 / 002562 Al entitled “Image Sensor with a Plurality of Super Pixels”, there is no 1 to 1 correspondence between event pixels and frame pixels, but it is on the same sensor.
[0082] Thus, in a more general way, the system can combine two separate sensors with different spatial resolutions. Moreover, a typical conventional sensor uses a Bayer color filter to produce a color image. Likewise, it is possible to put a Bayer color filter on an event-based sensor. In this case, the different operations mentioned above (deblurring, temporal shift and re-blurring) need to be applied per color channel. However, most event-based sensors do not have a Bayer color filter and only record luminance changes. Thus, the different modules mentioned above can only be applied to grayscale images.
[0083] Thus, in order to be able to use frame-based and event-based sensors with different numbers of pixels (also called sensor resolution) and to take into account the differences in color space between frame and event sensors, the present invention proposes a resolution and color restoration module.
[0084] This module can refer to the mapping from the recorded event buffer to an event buffer ready to be used independently or in series for the deblurring, image shift and re-blurring modules. It can also perform this transformation at different stages of the method or at the output.
[0085] In addition, sensors are not ideal and more specifically have thresholds that can vary with different factors and are subject to noise, the output of the resolution and color restoration module is still expected to ideally follow the mathematical definition of “reference color” events. This resolution and color restoration module thus also acts as a noise and artifact removal algorithm.
[0086] In general, this module can be implemented in several ways to treat each of the resolution, color restoration and idealization problems:
[0087] • Use methods implemented in a cascaded and independent manner, such as reprojection algorithms, which can be based on calibrated distortion, internal and external parameters of two sensors, based on machine learning, or a combination of methods based on calibrated distortion, internal and external parameters with machine learning methods for optimization, piped with colorization algorithms that can be based on machine learning or model-based, piped with a noise removal module, or possibly based on machine learning.
[0088] • Use the machine learning methods from both of these modules and pair the resulting model with the remaining modules;
[0089] • Use an end-to-end machine learning approach to model the entire mapping.
[0090] These methods can be arranged in series in any order.
[0091] The following section provides some examples of different sub-modules that implement the resolution and color restoration modules.
[0092] -Event-based resolution and color restoration submodule
[0093] This submodule transforms the event buffer before applying the previously mentioned modules.
[0094] In practice, the frame and event sensors are exposed to the same scene. Therefore, it is possible to map the original buffer e = (e_i)_i = 1..N of events acquired during the exposure duration of the blurred frame to a buffer e' = (e'_i)_i = 1..N', so that the coordinates of event e'_i are in the coordinate system of the frame sensor. It should be noted that the set e' does not need to have the same number of events as e.
[0095] Furthermore, applying the image shifting module, image deblurring module, and image reblurring module to a color image can be accomplished by preprocessing the event stream to determine which change in the color channel triggered the change in luminance.
[0096] This preprocessing can be considered as, after selecting the 3D color space, finding the mapping of the event buffer e'=(e'_i)i=1..N', which is aligned with the frame-based coordinates, to the event buffer e”=(e”_i)_i=1..N” of the reference color, so that...
[0097] e”_i exists at pixel x in any of the following cases:
[0098] ·log(Ch(x,t) / Ch(x,t_r))>c+_ch and is then assigned polarity 1 and channel Ch
[0099] • log(Ch(x, t) / Ch(x, t_r)) < c-_ch and then assigned a polarity 0 and channel Ch
[0100] where Ch is any of the three color channels and c+_ch and c-_ch are color thresholds depending on the e' to e" mapping algorithm. An event in e" thus has additional color channel information. And in this case, for each event in e' we can produce several events or no event in e" depending on the mapping function used.
[0101] To go from the set e to the set e", in general, we need a model or a learned function f(e) -> [e"] that produces a list of output events from different color references given an input event.
[0102] By knowing the internal parameters of each camera (camera matrix K and lens distortion parameters D), the rotation R and translation T between the two cameras and the depth map of the scene, we can compute a map M that gives for each pixel of the frame camera the coordinates on the event camera. This map can then be used to transform the coordinates of each event to bring it to the coordinates of the frame camera. Then we can represent an event as a vector v = [x, y, t, p] where (x, y) is the position of the event, t is the time and p is the polarity; or as a binary vector b = [0...0 1 0...0] or a float vector v^f = [f_1, f_2, f_3, f_4] that encodes v in a different space. Then f(e) can be a linear or non-linear function of v, b or v^f that results in a matrix in the output where each column is a vector [x, y, t, p, ch] where ch is the color channel. The matrix can have additional rows that represent whether the column is a valid event or not. This function can be a concatenation of linear operators where the warping functions of the linear operators are obtained from M or can be learned given enough instances of input / output pairs.
[0103] - Intermediate resolution and color restoration submodule
[0104] Another possible example is by acting on the intermediate maps used by the above modules. First, consider that the temporal shift, deblur and reblur modules all act on quantities h(E[t_i, t_j], c-, c+) that can be discretized in time bins for computational efficiency.
[0105] The input to the temporal shift, deblur and reblur modules can be the sum of the binned events. For example, using binning with a fixed time bin width ΔΤ, one can use the quantities (h(E[t_0+i*ΔΤ, t_0+(i+1)*ΔΤ], c-, c+)_i=1..B as input. One can also use the quantities (∑_i=0..B h(E[t_0+i*ΔΤ, t_0+(i+1)*ΔΤ], c-, c+)_i=1..B as input. e+∈i*Δk..(i+1)*Δk,ep>0 c ++∑ e-∈i*Δk..(i+1)*Δk,ep<0 c - )_i=1..Bto complete the discretization with a fixed number of events, where Δk is the bin width in terms of number of events.
[0106] This intermediate resolution and color restoration submodule can be implemented by taking this into account and thus having its constitutive modules' inputs and outputs, whether said modules are run separately or abstracted together via a machine learning pipeline. In practice, the quantity h(E[t_I,t_j],c-,c+) can be represented as a matrix H in R m_eb×n_eb×B with m_eb×n_eb the resolution of the event-based sensor and B the number of time bins with fixed time width or with fixed number of events. Thus, this module can be defined as a function f: R m_eb×n_eb×B-> R m_fb×n_fb×B'×ch , transforming the equivalent matrix with the same frame dimensions and a selected number of bins B'. And in this case, using the calibration of the camera, f can contain the warping operator obtained from the camera calibration, or it can be a guided or joint bilateral filter, or it can learn the input / output pairs.
[0107] - Post-resolution and color restoration submodule
[0108] Finally, the above modules can be applied by converting the frame F to the resolution of the events and removing the color information, then deblurring this version of the frame represented with F d , and then transforming the deblurred F d represented with D d to the original frame resolution and reapplying the color information.
[0109] In this case, the geometric mapping M that maps two camera pixels is applied to deblur F d . It is also possible to keep the original resolution of the frame camera and apply some simple interpolation to the quantity h(E[t_I,t_j],c-,c+). Once we have D d , we need to transform it into a function f that is a color version (and with the right details at the resolution of the frame camera). This function can be a guided or joint bilateral filter, or it can learn the input / output pairs.
[0110] • Non-linear compensation module
[0111] The above module can be applied to the images acquired from the frame-based sensor before or after any non-linear processing done by the ISP. In the latter case, it is necessary to compensate also the non-linear processing in the events. Thus, the non-linear compensation module should be defined as follows.
[0112] Assuming that the camera response function is the identity, then equations (2), (3) and (4) describing the image time shift, deblur and reblur modules can be written as
[0113] F_1 = F_2 * f((e_i)_i=1..N) where F_1, F_2 are frames and (e_i)_i=1..N is the event stream and f is the transformation function.
[0114] Thus, assuming that the frame of interest is F_2, we can retrieve said frame by simple division.
[0115] If the camera response function is some function g, then the above relation becomes
[0116] g(F_1) = g(F_2 * f((e_i)_i=1..N))
[0117] Non-linear processing usually includes gamma correction. In this setting, we get
[0118] (F_1)^γ = (F_2)^γ * f((e_I)_I=1..N)^γ
[0119] Retrieval of the frame of interest is obtained by (F_2)^γ = (F_1)^γ / f((e_I)_I=1..N)^γ.
[0120] A non-linear compensation module applied to some function f of the event buffer (e_i)_i=1..N is defined to produce the value f((e_i)_i=1..N)^γ in order to compensate for the non-linearity in the output function of the frame-based sensor and to be able to apply the frame processing modules mentioned above. Thus, in general, in this way we can process any function g that can be factorized as the above gamma correction. However, when we apply the reblur module and in the case that the function g is not factorizable, this approach cannot be used. In this case, if g is invertible, we can compute g -1 (F_1) = F_2 * f((e_i)_i=1..N) and then when we obtain F_2, apply g(F_2) to the obtained F_2 with the same response function as F_1.
[0121] In order to further explain the invention, in the following some embodiments of the exemplary method according to the invention are described, where these embodiments use the above mentioned modules to enhance the image quality of an imaging system that uses an event stream to modify the captured image when capturing images in different ideal environments.
[0122] Figure 1Exposure windows for each row in a rolling shutter sensor are shown within a frame. Each black element represents the exposure window for a given row, and from one row to another, the exposure windows are shifted due to the rolling shutter, as they are read sequentially by the sensor circuitry. It should be noted that the exposure windows have all the same exposure duration or exposure time, but they start and end at different times for each sensor row. The light gray vertical bars represent the exposure windows for a global shutter sensor at the target time that the present invention intends to restore, where all the row exposure windows are aligned. In Figure 1 In the current embodiment, the rolling shutter sensor scans from the top (first) row to the bottom (last) row, so that a temporally distorted image (image with artifacts) is captured.
[0123] For each row as shown in Figure 1 , the time t_s^r inside the exposure window is marked as a reference. Alternatively, for a global shutter sensor, the exposure window starts and ends at the same time for all rows. In addition, the target time t_ref is also defined with a new target exposure duration (represented in gray bar shape).
[0124] In the following paragraphs, we will illustrate how to transform / correct the image frame as represented by the black elements, captured by the exposure windows of a rolling shutter sensor, into another one as represented by the gray bar shape, as shown in Figure 1 , where for each row r of the rolling shutter sensor, we have the start of the exposure duration se^r, the end of the exposure duration ee^r, and the reference time t_s^r in [se^r, ee^r]. The target exposure duration T_ref and the target time t_ref are also defined together.
[0125] Figure 2 A first exemplary method according to the present invention is shown, which comprises the following steps, where an image frame B is captured by sensor element rows of a rolling shutter sensor in a frame-based camera, each row exposure window of the frame-based camera is shown as a black element in Figure 1 , and an event stream E = [e_1, e_n] corresponding to the scene depicted in the image frame B is produced by an event-based camera, and where the frame-based camera and the event-based camera have different resolutions and color spaces.
[0126] Step S1 : Invert the response function g -1 (B) by applying B' = g
[0127] Step S2: For each row r of B':
[0128] a. Apply an image deblurring module. This corresponds to:
[0129] i. Step S2.1 : estimate the corresponding row r e in the event stream and compute the row d r e (t s r ) from the event stream, d r e (t s r ) = 1 / T ∫ se^r ee^r e h(E^r^e[t_s^r,t],c+,c-) dt.
[0130] ii. Step S2.2: if the event stream is at a different resolution and color than the image, apply the intermediate resolution and color restoration submodule to obtain the row d r (t s r ) corresponding to row r of B.
[0131] iii. Step S2.3: finally compute the sharp row at t s r, for L r (t s r ) = B r / d r (t s r ).
[0132] b. For each time t d in [t ref - T ref / 2, t ref + T ref / 2], apply the image time shift module. This corresponds to:
[0133] i. Step S2.4: estimate the corresponding row r e in the event stream and compute the row e r e (t d ) from the event stream, e r e (t d ) = e h(x ,E[t_s^r,t_d],c+,c-) .
[0134] ii. Step S2.5: if the event stream is at a different resolution and color than the image, apply the intermediate resolution and color restoration submodule to obtain the row e r (t d ) corresponding to row r of B.
[0135] c. Step S2.6: apply the image de-blurring module and estimate the row r of the image at time t ref, i.e. L r (t ref ) = L r (t s r ) / T ∫ t_ref+T_ref / 2 t_ref-T_ref / 2 e r (t)dt.
[0136] These steps S2.1 to S2.6 in step 2 will be repeated until the last row of B is processed, and then move to next step 3.
[0137] Step S3: apply g to L and obtain the final corrected image L', similar to the image that would have been captured by a global shutter sensor at target time t ref with target exposure duration T ref.
[0138] With these steps, it is possible to correct an image frame obtained by a frame-based camera to a correct image using an event stream.
[0139] In the above embodiments, the term "sensor row" is used only for presentation clarity. It should be understood that the above equations are applied pixel by pixel in each row. Likewise, due to different lens distortions, different sensor sizes and orientations, and distances of objects in the scene, a row in the frame-based sensor does not necessarily correspond to a row in the event-based sensor, but to one or several non-continuous curves in the event-based sensor. The term "estimating the corresponding row in the event stream" means performing the operation for each pixel in the event-based sensor that lies on the curve(s) corresponding to the row in the frame-based sensor. Thus, each processing of the event data produced by the pixels in a row of the event-based sensor can use a different t s r depending on the frame-based sensor row it pertains to.
[0140] Alternatively, the exposure duration T ref can be short enough so that there is no need to apply the image deblurring module. In this case, there is no step c (step S2.6). Having a large target exposure duration T ref has only the effect of adding some artistic blur in the resulting image.
[0141] Figure 3 A second exemplary method according to the application is shown, which differs from the first exemplary method in that no deblurring module is used and that the camera response function g is factorizable, and thus the non-linear response function can be corrected without inverting g. This second exemplary method comprises the following steps:
[0142] Step S2: For each row r of B:
[0143] Step S2.1 : Estimate the corresponding row r e in the event stream, and compute the row d r e (t s r ) = 1 / T∫ se^r ee^r e h(E^r^e[t_s^r,t],c+,c-) dt.
[0144] Step S2.2: If the event stream is at a different resolution and color than the image, apply an intermediate resolution and color restoration submodule to obtain a row d r (t s r ) corresponding to the row r of B.
[0145] Step S2.2b: Correct d r (t s r ) with a non-linear correction module.
[0146] Step S2.3: Then finally compute the sharp row at t s r, for L r (t s r ) = B r / d r (t s r ).
[0147] These steps in this second exemplary method will be repeated until the last row of B is processed.
[0148] In the above exemplary method, intermediate resolution and color restoration sub-modules are used. However, compared to the exemplary embodiments as shown in Figure 2 In order to use event-based resolution and color correction sub-modules, in a third exemplary method as shown in Figure 4 In the third exemplary method as shown in Figure 4 In order to use event-based resolution and color correction sub-modules, in a third exemplary method as shown in
[0149] As another variant, Figure 5 A fourth exemplary method according to the present application is shown, wherein post resolution and color correction sub-modules are used. In this case, compared to the exemplary embodiments as shown in Figure 2 In order to use event-based resolution and color correction sub-modules, in a third exemplary method as shown in Figure 5 In order to use event-based resolution and color correction sub-modules, in a third exemplary method as shown in
[0150] In addition, the present application also proposes some embodiments of frame enhancement with fully data-driven methods. In particular, the present application proposes deblurring, creating multiple frames from a single blurred frame, or enhancing the dynamic range of frames in an end-to-end manner with a machine learning model. Using the additional information provided by event-based sensors can reduce the model size and thus the amount of computation.
[0151] That is, the present application proposes to model, using machine learning:
[0152] • the transformation of a blurred frame B and the event stream (e_i)_i=1..N into a sharp image S at a given time inside or outside the exposure duration of the blurred frame, inside or outside the exposure duration thereof.
[0153] • transforming the blur frame B and the stream of events (e_i)_i=1..N during its exposure duration into a set of N clear frames (S_0,..., S_N) regularly spaced in time during the exposure duration of the blur frame.
[0154] • transforming the frame F1 together with the events (e_i)_i=1..N during the exposure duration or any pre-processing of this stream of events into a frame F2 with enhanced dynamic range.
[0155] The present invention also proposes variations of the above mapping, where the stream of events (e_i)_i=1..N is replaced by any sufficient pre-processing.
[0156] In addition, by means of the above modules, sub-modules and methods according to the present invention, it is possible, in some other embodiments, to obtain slow motion videos and videos with enhanced dynamic range, which will be described hereafter.
[0157] For slow motion videos, the exemplary method described above can produce artifact-free images at a specific time t_ref. However, it is also possible to produce more than one corrected image in the output for each image in the input. Thus, embodiments according to the present invention can produce several frames in the output and increase the frame rate of the produced video, as shown in the following Figure 6
[0158] Figure 6 Exposure windows for two consecutive images captured by a frame-based camera are shown. Similar to one of Figure 1 , which consists of tilted black elements representing the exposure window of each line rolling shutter, and the exposure window is transformed into a plurality of corrected exposure windows, i.e. light grey vertical bars, at 13 different times, representing 13 frames at these different times, so that the frame rate is greatly increased.
[0159] Thus, it is possible to produce high frame rate videos without the limitation on the exposure duration that occurs when shooting high frame rate videos with frame cameras only (because the exposure duration must be smaller than 1 / (frames per second)). By using the present invention, it is possible to first shoot a video of e.g. 60 frames / second (thus with a maximum exposure of 16 ms), and then to produce a video with a different (typically higher) frame rate per second, where the maximum frame rate per second is only limited by the temporal resolution of the event camera.
[0160] For videos with enhanced dynamic range, in addition to the above embodiments aiming at enhancing the production of a frame with only another frame in the input and the stream of events, it is also possible to align and merge several images (or frames) captured sequentially by the sensor. Using several images corrected / enhanced by the present invention allows improving the quality of the image or video and extending its dynamic range.
[0161] This improves over the state of the art where several frames are acquired and spatially aligned before being averaged to reduce image noise. It is difficult to obtain alignment when the scene changes during frame acquisition, due to moving objects in the scene or due to relative motion between the scene and the camera. State of the art approaches typically use computationally intensive motion estimation and motion compensation to align the frames. In the present invention, it is possible to use the event stream and the image time shift module to "align" several frames to a unique t_ref in a more efficient way without explicit motion compensation.
[0162] The present invention is particularly advantageous when the exposure duration between frames acquisition changes. Changing the exposure duration is typically used to better capture the dynamic range of a scene: if the highest and lowest luminosity of a scene exceed the highest and lowest light intensity that can be represented in a single exposure due to the physical limitations of the (frame-based) sensor, it is common to capture N frames with different exposure durations E_i, i in [0,N), and to merge the frames into one "composite" frame, thereby merging both the highest and lowest light intensity into a single frame. The varying exposure of the N frames makes the alignment more difficult for motion compensation techniques in the state of the art.
[0163] In this case, the above mentioned modules and methods can be used to capture several images of a scene and align them to a unique target time t_ref in order to obtain a High Dynamic Range (HDR) result.
[0164] Thus, the present invention proposes to use a combination of a traditional frame-based image sensor and an event-based sensor, such as a DVS sensor, to produce image / video frames and event streams, and then to combine the image / video frames with the event data in order to more efficiently remove artifacts and enhance the dynamic range.
[0165] Furthermore, the aforementioned modules / methods according to the present invention can be implemented in many ways, such as program instructions for execution by a processor, as a software module, as microcode, as a computer program product on a computer readable medium, as a logic circuit, as an application specific integrated circuit, as firmware, etc. Embodiments of the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment containing both hardware and software elements. In a preferred embodiment, the present invention is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.
[0166] Furthermore, embodiments of the present application can take the form of a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer, processing apparatus, or any instruction execution system. For the purposes of this description, a computer-usable or computer readable medium can be any apparatus that can contain, store, communicate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The medium can be electronic, magnetic, optical, or semiconductor system (or apparatus or device). Examples of a computer-readable medium include but are not limited to a semiconductor or solid state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a rigid magnetic disk, an optical disk, or the like. Current examples of optical disks include compact disk - read only memory (CD-ROM), compact disk - read / write (CD-R / W) and DVD.
[0167] The embodiments described above are intended to be illustrative only. Numerous modifications can be made by those skilled in the art without departing from the true spirit and scope of the application as defined by the following claims.
Claims
1. Image enhancement method comprising: - capturing a reference image frame of a scene by a frame-based rolling shutter sensor containing rows of sensor elements, wherein the reference image frame contains artifacts and has an exposure duration; - obtaining, at least during the exposure duration, an event stream by an event-based sensor synchronized with the frame-based sensor, wherein the events encode luminance variations of the scene corresponding to the reference image; and - deriving a corrected image frame, free of the artifacts, by processing each row of the reference image frame using the event stream.
2. The image enhancement method of claim 1, wherein, The reference image frame is obtained at a reference time (t_r) and the step of deriving a corrected image frame comprises a step of image temporal shift, wherein the reference image frame is shifted to the corrected image frame at a target time (t) inside or outside the exposure duration by establishing a relationship between the events and the reference image frame and accumulating luminance variations encoded by events obtained during a time interval [t_r, t] between the reference time and the target time.
3. The image enhancement method of claim 1, wherein, The reference image frame is a blurred frame and the method comprises a step of image deblurring for turning the blurred frame into a sharp frame.
4. The image enhancement method of claim 1, wherein, The method comprises a step of re-blurring the corrected image frame.
5. The image enhancement method of claim 1, wherein, The method comprises a step of non-linear compensation for compensating non-linear processing in the image frame and in the events.
6. The image enhancement method of any one of claims 1 to 5, characterized in that, The reference image frame has a first resolution and / or a first color space and the event stream has a second resolution and / or a second color space different from the first resolution and / or the first color space, and wherein the method further comprises a step of resolution and color recovery for compensating resolution and / or color space differences between the reference image frame and the event stream.
7. The image enhancement method of claim 1, wherein, The method comprises producing a plurality of corrected image frames at different time distributions inside or outside the exposure duration.
8. The image enhancement method of claim 1, wherein, The method comprises aligning a plurality of corrected image frames derived in different exposure durations.
9. Image enhancement system comprising: - a frame-based rolling shutter sensor containing rows of sensor elements configured to obtain a reference image containing artifacts and having an exposure duration, - an event-based sensor synchronized with the frame-based sensor configured to obtain, at least during the exposure duration, an event stream encoding luminance variations of the scene corresponding to the reference image; - an image enhancement module configured to derive from the reference image a corrected image frame, free of the artifacts, using the event stream.
10. The image intensifier system of claim 9, wherein, The reference image is obtained at a reference time and the image enhancement module comprises an image temporal shift module, wherein the reference image frame is shifted to the corrected image frame at a target time inside or outside the exposure duration by establishing a relationship between the events and the reference image frame and accumulating luminance variations encoded by events obtained during a time interval between the reference time and the target time.
11. The image intensifier system of claim 10, wherein, The image enhancement module comprises a non-linear compensation module for compensating non-linear processing in the image frame and in the events.
12. The image intensifier system of claim 10, wherein, The reference image has a first resolution and / or a first color space, and the event stream has a second resolution and / or a second color space different from the first resolution and / or the first color space, and wherein the image enhancement module further comprises a resolution and color restoration module for compensating for resolution and / or color space differences between the reference image and the event stream.
Citation Information
Patent Citations
Real-time video deblurring
EP2454876B1
Systems and methods for rolling shutter compensation using iterative process
US10194101B1
Methods and apparatus for image deblurring and sharpening using local patch self-similarity
US8687913B2
In-stream rolling shutter compensation
US9554048B2
Methods and systems for removal of rolling shutter effects
WO2014042894A1