Image enhancement method, image enhancement device, and image enhancement system

The combination of frame-based and event-based sensors in image processing techniques addresses motion artifacts and dynamic range challenges, achieving high-quality image and video capture in low-light conditions with reduced computational demands.

JP7767331B2Active Publication Date: 2025-11-11プロフジー
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
JP2022579822
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-23
Filing Date
2021-06-23
Publication Date
2025-11-11
Estimated Expiration
2041-06-23

AI Technical Summary

Technical Problem

Conventional image and video cameras face limitations such as motion artifacts due to rolling shutter readout, exposure duration tradeoffs leading to motion blur or sensor noise, and challenges in capturing high dynamic range scenes with computational complexity.

Method used

An image enhancement method combining frame-based and event-based sensors to synchronize and process image frames and event data, using modules for image time shifting, deblurring, reblurring, and resolution/color restoration to correct artifacts and enhance dynamic range.

Benefits of technology

Enables high dynamic range image and video capture of low-light, motion-filled scenes with reduced computational complexity, effectively removing artifacts and enhancing image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007767331000001
    Figure 0007767331000001
  • Figure 0007767331000002
    Figure 0007767331000002
  • Figure 0007767331000003
    Figure 0007767331000003
Patent Text Reader

Abstract

- acquiring, by a frame-based sensor, a reference image frame of the scene, the reference image including the artifact and having an exposure duration; - acquiring a stream of events by an event-based sensor synchronized with the frame-based sensor for at least the duration of the exposure, the events encoding brightness changes of the scene corresponding to the reference image; - deriving artifact-free corrected image frames from reference image frames using the stream of events; An image enhancement method comprising:
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to image processing techniques, and in particular to image enhancement. [Background technology]

[0002] Conventional image and video cameras and systems are generally limited by several factors. Specifically, at the sensor design level, in order to create smaller image sensors with lower cost and greater sensitivity, image sensors often use rolling shutter readout. While this readout offers advantages, it can also lead to artifacts due to motion occurring between frame captures.

[0003] In addition, when videos and images are captured by a conventional camera, camera settings such as changing the exposure duration result in a different ISO tradeoff for capture, like the normal exposure / ISO tradeoff, where more or less light is integrated, but there is still a risk of creating a motion-blurred image. Maintaining a short exposure duration can reduce motion blur, but in low-light situations, this will create an image that is too dark. In the latter case, it is common to compensate for the smaller input signal by increasing the ISO gain of the sensor. However, this also amplifies the sensor noise.

[0004] Furthermore, some scenes may contain both bright and dark areas. For these scenes, one must often choose between an exposure that exposes the bright areas well and another that exposes the dark areas better. A typical solution to this problem is to combine several images and compress the measured light intensities using some nonlinear function to create an HDR image. However, this solution requires the alignment of several images, typically two to eight, and is therefore computationally too complex for high-frame-rate video for traditional image and video capture.

[0005] In the art, there are already some solutions that attempt to compensate for the above-mentioned limitations of conventional imaging and video capture. For example, WO2014042894A1, U.S. Patent No. 10194101B1, and U.S. Patent No. 9554048B2 disclose methods for removing certain rolling shutter effects, WO2016107961A1, EP2454876B1, and U.S. Patent No. 8687913B2 disclose methods for removing blur without the rolling shutter effect, and some other papers discuss computing high dynamic range (HDR) images and videos without removing rolling shutter and blur. However, they are usually tailored to certain artifacts and / or involve significant computational complexity. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] WO2014042894A1 [Patent Document 2] U.S. Patent No. 10194101B1 [Patent Document 3] U.S. Patent No. 9554048B2 [Patent Document 4] WO2016107961A1 [Patent Document 5] EP2454876B1 [Patent Document 6] U.S. Patent No. 8687913B2 [Patent Document 7] WO2020 / 002562A1 Summary of the Invention [Problem to be solved by the invention]

[0007] The object of the present invention is to remedy these weaknesses. [Means for solving the problem]

[0008] In this regard, in accordance with one aspect of the present invention, there is provided an image enhancement method, the method comprising: - acquiring, by a frame-based sensor, a reference image frame of the scene, the reference image frame including the artifact and having an exposure duration; - acquiring a stream of events by an event-based sensor synchronized with the frame-based sensor for at least the exposure duration of a reference image frame, the events encoding luminance changes of a scene corresponding to the reference image; - deriving artifact-free corrected image frames from the reference images using the stream of events; Includes:

[0009] Thus, such an arrangement according to the present invention combines frame images and event data to increase dynamic range and remove image artifacts from the frame images, such as wobble (also known as the Jello effect), skew, spatial aliasing, and temporal aliasing, particularly for scenes with motion in low light.

[0010] In one embodiment according to the present invention, a reference image frame is acquired at a reference time, and the step of deriving a corrected image frame includes a step of image time shifting, in which the reference image frame is shifted to a corrected image frame at a target time within or other than the exposure duration by establishing a relationship between the event and the reference image frame and accumulating brightness changes encoded by the event acquired during the time interval [t_r,t] between the reference time and the target time to shift the image captured at a different time, i.e., at the target time.

[0011] Furthermore, according to the present invention, the reference image frame is a blurred frame, and the method includes a step of image blur correction to restore the blurred frame to a clear image, for example, to obtain a clear reference image when the exposure duration is too long.

[0012] Alternatively, the method according to the invention comprises the step of re-blurring the corrected image frame so as to simulate from the sharp image a larger exposure duration effect which makes the image more realistic, especially in fast motion scenes.

[0013] Alternatively, the method according to the invention includes a step of non-linear compensation to compensate for non-linear processing in both the image frames and the events.

[0014] Furthermore, in another embodiment of the present invention, the reference image has a first resolution and / or a first color space and the stream of events has a second resolution and / or a second color space different from the first resolution and / or first color space, and the method further comprises a resolution and color restoration step to compensate for the difference in resolution and / or color space between the reference image and the stream of events.

[0015] In yet another embodiment of the present invention, the image frames are generated by a frame-based camera including a rolling shutter sensor including rows of sensor elements (pixels), and for example, the step of deriving the corrected image frames proceeds sequentially for each row of the rolling shutter sensor.

[0016] In yet another embodiment of the present invention, a method includes generating a plurality of corrected image frames distributed at different times in the exposure duration.

[0017] Alternatively, the method according to the invention comprises the step of registering a plurality of corrected image frames derived at different exposure durations.

[0018] According to another aspect of the present invention, there is further provided an image enhancement apparatus, the image enhancement apparatus comprising: - a frame-based camera including a frame-based sensor; an event-based camera including an event-based sensor, wherein the frame-based camera and the event-based camera observe the same scene and each generate an image frame and a stream of events corresponding to the scene, and the frame-based camera and the event-based camera are synchronized; and - a processor for carrying out the method described above Equipped with.

[0019] Alternatively, according to yet another aspect of the present invention, there is further provided another image enhancement device, the image enhancement device comprising: a camera adapted to generate image frames and a stream of events corresponding to a scene; - a processor for carrying out the method described above Equipped with.

[0020] Furthermore, according to yet another aspect of the present invention, there is further provided an image enhancement system, the image enhancement system comprising: a capture module configured to acquire a reference image, the reference image having an exposure duration, the reference image possibly including artifacts, and to acquire a stream of events corresponding to the scene during the exposure duration of the reference image frame; an image enhancement module adapted to derive artifact-free corrected image frames from a reference image using the stream of events; Equipped with.

[0021] In one embodiment of the system according to the present invention, the image enhancement module includes an image time shifting module for shifting the reference image frame to the corrected image frame at any time (t) within or outside the exposure duration by accumulating the event-encoded luminance changes.

[0022] In another embodiment of the system according to the present invention, the image enhancement module comprises a non-linear compensation module for compensating for non-linear processing in the image frames and events.

[0023] In yet another embodiment of the system according to the invention, the reference image has a first resolution and / or a first color space and the stream of events has a second resolution and / or a second color space different from the first resolution and / or first color space, and the image enhancement module further comprises a resolution and color restoration module for compensating for differences in resolution and / or color space between the reference image and the stream of events.

[0024] The present invention extends the capabilities of standard frame-based image sensors by pairing them with an event stream, for example, from a DVS sensor, extending the traditional tradeoffs of image sensors and enabling video and image capture that better matches scene characteristics. In this regard, the present invention enables high dynamic range image and video capture of low light, motion-filled scenes with limited computational complexity requirements.

[0025] Other features and advantages of the present invention will be seen in the following description, taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0026] [Figure 1] 1 shows exposure windows for an image captured by a frame-based sensor and an image corrected by a method according to the present invention; [Figure 2] FIG. 2 is a functional block diagram of a first exemplary method according to the present invention. [Figure 3] FIG. 4 is a functional block diagram of a second exemplary method according to the present invention. [Figure 4] FIG. 10 is a functional block diagram of a third exemplary method according to the present invention. [Figure 5] FIG. 10 is a functional block diagram of a fourth exemplary method according to the present invention. [Figure 6] FIG. 1 shows two images captured by a frame-based sensor and exposure windows for multiple images corrected by a method according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0027] The context of the present invention includes an exemplary image capture architecture comprising a conventional image frame-based camera and an event-based camera, where the frame-based camera and the event-based camera are observing the same scene where objects or the entire scene are moving relative to the frame-based camera and the event-based camera, and the frame-based camera captures and generates image frames of the scene (including the moving objects), and the event-based camera generates a stream of events corresponding to the same scene including the moving objects captured during the capture of the image frames (frame period), and possibly before and after.

[0028] Specifically, a frame-based camera has a rolling shutter sensor having rows of sensors, such as CMOS sensors, that are scanned row by row sequentially across a scene at high speed so that an image is captured or generated as the sensor scans from the first row to the last row in one period called a frame period.

[0029] An event-based camera includes an asynchronous vision sensor facing the same scene observed by the frame-based camera and receiving light from the same scene. The asynchronous vision sensor includes an array of sensing elements (pixels), each containing a light-sensitive element. Each sensing element (pixel) creates an event corresponding to a variation in light in the scene. The asynchronous vision sensor can be a contrast detection (CD) sensor or a dynamic vision sensor (DVS).

[0030] Alternatively, the present invention may use another possible exemplary image capture architecture that includes a single sensor that records both frames and events, such as a DAVIS sensor or an asynchronous time-based image sensor (ATIS).

[0031] In the present invention, it is necessary to synchronize the frame data with the event data.

[0032] Specifically, in the case of a combination of an image frame-based sensor and a DVS or CD sensor, the time bases of the two sensors can be synchronized by exchanging synchronization signals between the sensors.

[0033] For example, one possibility is to rely on a flash-out signal from a frame-based camera that is typically used to trigger a flash during the exposure duration of a frame capture, and use this signal to signal the start of the exposure duration to the CD sensor. The CD sensor can then timestamp the received signal to its own time base, allowing the processing element to relate the event timestamp to the exposure duration of the frame-based sensor.

[0034] Such synchronization is important because the present invention is based on the relationship between frames and events, which relationships are discussed hereinafter.

[0035] The events of an event-based camera (DVS or CD) are related to the temporal evolution of luminance in a scene. The event-based sensor records "events" that encode changes in luminance relative to a reference state, called the reference image L(t_r) at a reference time t_r. The event-based camera then: log(L(x,t) / L(x,t_r))>c+ then create a positive event at time t at (pixel) location x, log(L(x,t) / L(x,t_r)) <c- If so, it will create a negative event.

[0036] where c+ and c- are the positive and negative thresholds, respectively. Note that the use of logarithmic functions in this case is merely an example, and these relationships can be applied to any monotonic and continuous functions. Also, the polarity associated with a particular event e is determined by the p is defined by e p teeth, e p =τ(log(L(x,t) / L(x,t_r)),c+,c-) where τ(x,c+,c-)=1 if x>=c+, 0 if x∈(-c-,c+), and -1 if x≦-c-.

[0037] When a series of events E[t_r, t_n] = e1, …, ek are received at position x within the time interval T = [t_r, t_n], Σ e∈E[t_r,t_n] e p = τ(log(L(x, t1) / L(x, t_r)), c+, c-) + … + τ(log(L(x, t_n) / L(x, t_(n - 1))), c+, c-) ≒ log(L(x, t_n) / L(x, t_r)), c+, c- is obtained.

[0038] h(x, E[t_r, t_n], c+, c-) = Σ e+∈E[t_r,t_n],ep>0 c++Σ e-∈E[t_r,t_n],ep<0 c- By defining (1) h(x, E[t_r, t_n], c+, c-) ≒ log(L(x, t_n) / L(x, t_r), c+, c-) is obtained.

[0039] Equation (1) is applied independently to each pixel.

[0040] In this case, Equation (1) establishes the relationship between two images at different times and the events received between these times. Next, using this relationship, a method for manipulating and improving images according to the present invention will be discussed.

[0041] In the context of the present invention, the present invention also comprises a processor or computer calculation unit adapted to execute an image enhancement method based on the following modules.

[0042] Image time shifting module From Equation (1), by accumulating the luminance changes encoded by the events E[t_r, t] and updating the luminance L(x, t_r), a sharp image L(x, t) at any time t from the sharp reference frame L(x, t_r) when t_r < t can be obtained (2) L(x, t) ≒ L(x, t_r)e h(x,E[t_r,t],c+,c-) It can be derived as:

[0043] If t_r>t, h(x,E[t_r,t],c+,c-)=-(Σ e+∈E[t_r,t],ep>0 c++Σ e-∈E[t_r,t],ep<0 c-) is.

[0044] Equation (2) allows us to estimate the image at any time t given a reference image at time t_r and the events recorded in the time interval [t_r,t]. This equation forms one of the basic modules of the method according to the invention, the image time shifting module, which allows us to "shift" images captured at different times.

[0045] Image Deblur Module Equation (2) assumes that a reference image is available at time t_r. This may be true if the exposure duration is short enough (i.e., the image is sharp). However, if the exposure duration is too long, or if there is motion in the scene, or relative motion between the scene and the camera, a sharp reference image may not be available.

[0046] Therefore, we need to ensure that we shift the sharp image at t_r. Mathematically, the blur frame is B=1 / T∫ f+T / 2 f-T / 2 L(t)dt where B is the blurred image, f is a given time within the exposure duration and is used as the reference time, T is the exposure duration, and L(t) is the sharp image at time t.

[0047] Substituting equation (2) into place of L(t), (3) L(f)≒T*B / ∫ f+T / 2 f-T / 2 e h(E[f,t],c+,c-) dt can be obtained.

[0048] Equation (3) defines the image deblurring module, which allows to recover a sharp image from a blurred image at any arbitrary time within the exposure duration. The image deblurring module and the image time shifting module allow to reconstruct a sharp image at any desired time t.

[0049] Image Reblur Module Photographers or video makers sometimes want to maintain / add some blur to their images / videos. In this regard, an image can be simulated with any exposure duration T_n by simply averaging all the possible different images that occurred during the exposure. For example, if a sharp image is reconstructed at time t_r and exposure T_n centers the image at time t_r, then it is (4) B≒L(t_r) / T∫ t_r+T_n / 2 t_r-T_n / 2 e h(E[t_r,t],c+,c-) dt This can be computed by:

[0050] Equation (4) defines the image reblur module, where a longer exposure duration is simulated from the sharp image.

[0051] Resolution and Color Restoration Module Each of the modules defined above works independently for each pixel. The modules also assume that there is a one-to-one correspondence between event-based sensor pixels, such as DVS sensors, and frame pixels, and that c+ and c− are known for each pixel. However, this may not be the case for sensor combinations. For example, in WO 2020 / 002562 A1, entitled "Image Sensor with a Plurality of Super Pixels," event pixels and frame pixels are in the same sensor, but there is no one-to-one correspondence between them.

[0052] Therefore, in a more general way, the system can combine two separate sensors with different spatial resolutions. Moreover, typical conventional sensors use a Bayer filter to create color images. Similarly, an event-based sensor can be equipped with a Bayer filter. In this case, the different operations mentioned above (deblurring, time shifting, and reblurring) must be applied to each color channel. However, most event-based sensors do not have a Bayer filter and only record changes in luminance. Therefore, the different modules mentioned above can only be applied to grayscale images.

[0053] Therefore, in order to enable the use of frame-based and event-based sensors with different numbers of pixels (also known as sensor resolution) and to take into account the differences in color space between frame and event sensors, the present invention proposes a resolution and color restoration module.

[0054] Such a module can independently or sequentially indicate the mapping from the recorded buffer of events to buffers of events ready for use by the deblur, image shifting, and reblur modules, and it can also perform this transformation at different stages of the method or at the output.

[0055] Additionally, although sensors are not ideal, more specifically, have thresholds that can vary due to different factors and are subject to noise, the output of the resolution and color restoration module is still expected to ideally follow the mathematical definition of a "color-referenced" event, so this resolution and color restoration module also serves as a noise and artifact removal algorithm.

[0056] In general, this module can be implemented in several ways, with each of the resolution, color restoration, and idealization problems addressed by either: - using techniques that are in turn implemented independently, such as a reprojection algorithm, which may be based on the calibrated distortion, intrinsic and extrinsic parameters of the two sensors, that are machine learning based, or blending a technique based on the calibrated distortion, intrinsic and extrinsic parameters with a machine learning technique for refinement, pipelined with a colorization algorithm, which may be machine learning based or model based, and also possibly pipelined with a machine learning based denoising module; - using machine learning techniques in two of these modules and pairing the resulting model with the remaining modules; - Model the entire mapping using end-to-end machine learning techniques.

[0057] These techniques can be sequenced in any order.

[0058] Next, some examples of different sub-modules that implement the resolution and color restoration module are given.

[0059] - Event-based resolution and color restoration submodule This submodule transforms the event buffer before applying the previously mentioned modules.

[0060] In fact, the frame sensor and the event sensor are exposed to the same scene. Therefore, it is possible to map the original buffer e=(e_i)_i=1..N of events captured during the exposure duration of the blurred frame to a buffer e'=(e'_i)_i=1..N' such that the coordinates of event e'_i are in the coordinate system of the frame sensor. Note that the set e' does not need to have the same number of events as e.

[0061] Also, the application of the image shifting module, the image blur correction module, and the image re-blurring module to color images can be performed by preprocessing the event stream to determine which color channel changes are triggering the change in luminosity.

[0062] This preprocessing can be thought of as finding a mapping from the frame-based coordinates and the buffer of already-aligned events e'=(e'_i)i=1..N' to the buffer of color-based events e''=(e''_i)_i=1..N'' after selecting a 3D color space, whereby e''_i is - log(Ch(x,t) / Ch(x,t_r))>c + _ch, and then either when returning to polarity 1 and channel Ch, - log(Ch(x,t) / Ch(x,t_r))<c-_ch, and then when returning to polarity 0 and channel Ch exists at pixel x in either of them.

[0063] Ch is any one of the three color channels, and c + _ch and c-_ch are color thresholds that depend on the mapping algorithm from e' to e''. Therefore, the events in e'' have additional color channel information. Also, in this case, for each event in e', several events in e'' may or may not be generated according to the mapping function used.

[0064] Therefore, to go from the set e to the set e'', generally, a model that generates a list of events with different color references as output given the events in the input, or a function f(e)→[e''] needs to be learned.

[0065] By knowing the internal parameters of each camera (camera matrix K and lens distortion parameter D), the rotation R and translation T between the two cameras, and the depth map of the scene, a map M can be computed that gives, for each pixel of the frame camera, a coordinate on the event camera. This map can then be used to transform the coordinates of each event to bring it into the coordinates of the frame camera. An event can then be represented as a vector v = [x, y, t, p], where (x, y) is the event's position, t is time, and p is polarity, or as a binary vector b = [0...0 1 0...0], or as a floating-point vector v̂f = [f_1, f_2, f_3, f_4], encoding v in a different space. f(e) can then be a linear or nonlinear function of v, b, or v̂f that gives at output a matrix whose each column is a vector [x, y, t, p, ch], where ch is a color channel. The matrix can have an additional row that indicates whether the column is a valid event or not. This function can be a linear operator concatenation of warping functions obtained from M, or it can be learned given enough examples of input / output pairs.

[0066] - Intermediate resolution and color restoration sub-modules Another possible example is by operating on the intermediate maps used by the above modules. First, consider that the time shifting, deblurring, and reblurring modules all operate on the quantity h(E[t_i, t_j], c−, c+), which can be discretized in time bins for computational efficiency.

[0067] The input to the time shifting, deblurring, and reblurring modules can be a collection of binned sums of events. For example, using binning with a fixed time bin width ΔT, one can use as input the quantity (h(E[t_0+i*ΔT, t_0+(I+1)*ΔT], c-,c+)_I=1..B. The discretization can also be done by dividing the quantity (Σ_ e+∈i*Δk..(i+1)*Δk,ep>0 c + +Σ e-∈i*Δk..(i+1)*Δk,ep<0 c - )_i=1..B, this can be done with a fixed number of events.

[0068] The intermediate resolution and color restoration sub-modules can be implemented by taking this into account, and therefore having the inputs and outputs of its constituent modules, whether they operate separately or are abstracted together through a machine learning pipeline. In practice, the quantity h(E[t_I, t_j], c-, c+) is calculated using the matrix H\in R m_eb×n_eb×B where m_eb × n_eb is the resolution of the event-based sensor and B is the number of time bins, either with a fixed time width or a fixed number of events. This module then calculates a function f:R that transforms the matrix into an equivalent matrix with the same dimensions of the frame and the selected number of bins B'. m_eb×n_eb×B→ R m_fb×n_fb×B'×ch Also, in this case, using camera calibration, f can include a warping operator derived from camera calibration, or it can be a guided filter, or a joint bilateral filter, or it can be learned using input / output pairs.

[0069] - Post-resolution and color restoration sub-modules Finally, the above mentioned module converts the frame F to the resolution of the event, removes the color information, and then converts F d Then, blur correction is applied to such a version of the frame denoted by D dThe blur-corrected F is shown by d can be transformed to the original frame resolution and the color information reapplied.

[0070] In this case, the geometric mapping M that maps the two camera pixels is F d It is also possible to maintain the original resolution of the frame camera and apply some simple interpolation to the quantity h(E[t_I, t_j], c-, c+). Once D d Once we have , we need a function f (with the exact details at the resolution of the frame camera) that transforms it into a color version. This function can be a guided filter, or a joint bilateral filter, or it can be learned using input / output pairs.

[0071] - Nonlinear Compensation Module The above mentioned modules can be applied to the images captured from the frame-based sensor before or after any nonlinear processing performed by the ISP. In the latter case, it is necessary to compensate for the nonlinear processing even in the event. Therefore, the nonlinear compensation module shall be defined as follows:

[0072] Assuming the camera response functions are identical, equations (2), (3), and (4), which describe the image time shifting, deblurring, and reblurring modules, can be written as: F_1=F_2*f((e_i)_i=1..N) where F_1, F_2 are frames, (e_i)_i=1..N are event streams, and f is a transformation function.

[0073] So, assuming the frame of interest is F_2, it can be extracted by simple division.

[0074] If the camera response function is some function g, then the above relationship becomes g(F_1)=g(F_2*f((e_i)_i=1..N)) becomes.

[0075] Non-linear processing often includes gamma correction. In this setting: (F_1)^γ=(F_2)^γ*f((e_I)_I=1..N)^γ is obtained.

[0076] To extract the target frame, (F_2)^γ=(F_1)^γ / f((e_I)_I=1..N)^γ is obtained by

[0077] A nonlinear compensation module applied to some function f of a buffer of events (e_i)_i=1..N is defined to create a value f((e_i)_i=1..N)^γ to compensate for nonlinearities in the output function of the frame-based sensor and to allow the frame processing modules described above to be applied. So, in general, in this way we can work with any function g that can be factorized as the gamma correction described above. However, this approach cannot be used when applying a reblur module and if the function g cannot be factorized. In this case, if g is invertible, then g -1 We compute (F_1)=F_2*f((e_i)_i=1..N), and then once we have F_2, we can apply g(F_2), which has the same response function as F_1, to the resulting F_2.

[0078] To further explain the present invention, the present specification describes several embodiments of exemplary methods according to the present invention, which use the above-mentioned modules to enhance image quality of an imaging system that uses an event stream to modify captured images captured at different ideal settings.

[0079] FIG. 1 illustrates how rows in a rolling shutter sensor are exposed for a frame. Each black element represents the exposure window for a given row, which shifts due to the rolling shutter as it is sequentially read by the sensor circuitry from one row to the other. Note that the exposure windows all have the same exposure duration, or exposure time, but begin and end at different times for each sensor row. The light gray vertical bar represents the exposure window of the global shutter sensor at the target time that the present invention is intended to recover, where the exposure windows for all rows are aligned. In this embodiment of FIG. 1, the rolling shutter sensor scans from the top (first) row to the bottom (last) row, thereby capturing a temporally distorted image (an image with artifacts).

[0080] For each row shown in Figure 1, the time t_s^r within the exposure window is marked as a reference. Alternatively, for a global shutter sensor, the exposure window starts and ends at the same time for all rows. In addition, a target time t_ref is also defined by the new target exposure duration (shown by the gray vertical bar).

[0081] In the following paragraphs, we will clarify how an image frame captured by the exposure window of a rolling shutter sensor, represented by the black elements, can be transformed / corrected into another image frame, represented by the grey bars, by using the method according to the invention shown in Figure 1, where for each row r of the rolling shutter sensor, the start of the exposure duration se^r, the end of the exposure duration ee^r, and the reference time t_s^r\in[se^r, ee^r] are obtained. A target exposure duration T_ref is also defined together with the target time t_ref.

[0082] FIG. 2 shows a first exemplary method according to the present invention, comprising the following steps: an image frame B is captured by rows of sensor elements of a rolling shutter sensor in a frame-based camera, where the exposure window of each row is shown as a black element in FIG. 1; and a stream of events E=[e_1, e_n] corresponding to the scene depicted in image frame B is generated by an event-based camera, where the frame-based camera and the event-based camera have different resolutions and color spaces. Step S1: B' = g -1 Invert the response function g by applying (B). Step S2: For each row r of B', a. Apply the Image Deblur module, which corresponds to the following steps: i. Step S2.1: Estimate the corresponding row r^e in the event stream and calculate the row d^r^e(t_s^r)=1 / T∫ from the event stream. se^r ee^r e h(E^r^e[t_s^r,t],c+,c-) dt is calculated by computer. ii. Step S2.2: If the event stream is at different resolutions and colors of the image, apply the intermediate resolution and color restoration sub-module to obtain the row d^r(t_s^r) corresponding to the row r of B. iii. Step S2.3: Finally, compute the sharp lines in t_s^r as L^r(t_s^r)≈B'^r / d^r(t_s^r). b. For each time t_d\in[t_ref-T_ref / 2, t_ref+T_ref / 2], apply the image time shifting module, which corresponds to the following steps: i. Step S2.4: Estimate the corresponding row r^e in the event stream and calculate the row e^r^e(t_d)=e from the event stream. h(x,E[t_s^r,t_d],c+,c-) is calculated by computer. ii. Step S2.5: If the event stream is at different resolutions and colors of the image, apply the intermediate resolution and color restoration sub-module to obtain the row e^r(t_d) corresponding to the row r of B. c. Step S2.6: Apply an image reblur module to the image at time t_ref L^r(t_ref)=L^r(t_s^r) / T∫ t_ref+T_ref / 2 t_ref-T_ref / 2 Estimate row r of e^r(t)dt.

[0083] These steps S2.1 to S2.6 in step 2 are repeated until the last row of B is processed, and then the process moves to step 3. Step S3: Apply g to L to obtain a final corrected image L' similar to that captured by a global shutter sensor at target time t_ref with target exposure duration T_ref.

[0084] These steps allow the image frames acquired by the frame-based camera to be corrected using the stream of events.

[0085] In the above embodiments, the term sensor row is used merely for clarity of presentation. It should be understood that the above equation applies to each pixel in each row. Similarly, due to different lens distortions, different sensor sizes and orientations, and object distances in the scene, a row in the frame-based sensor does not necessarily correspond to a row in the event-based sensor, but may correspond to one or several non-contiguous curves in the event-based sensor. The term "estimating a corresponding row in the event stream" means performing an operation for each pixel in the event-based sensor that is on one or several curves corresponding to a row in the frame-based sensor. As a result, each processing of event data generated by pixels in a row of the event-based sensor can use a different t_s^r depending on the row of the frame-based sensor to which it pertains.

[0086] Alternatively, the exposure duration T_ref can be short enough so that there is no need to apply the image reblur module, in which case step c (step S2.6) does not exist. A long target exposure duration T_ref only has the effect of adding some artistic blur to the resulting image.

[0087] 3 illustrates a second exemplary method according to the present invention, which differs from the first exemplary method in that a reblur module is not used and the camera response function g is factorizable, so that nonlinear response functions can be corrected without inverting g. This second exemplary method includes the following steps: Step S2: For each row r of B, Step S2.1: Estimate the corresponding row r^e in the event stream and calculate the row d^r^e(t_s^r)=1 / T∫ from the event stream. se^r ee^r e h(E^r^e[t_s^r,t],c+,c-) dt is calculated by computer. Step S2.2: If the event stream is at different resolutions and colors of the image, apply the intermediate resolution and color restoration submodule to obtain the row d^r(t_s^r) corresponding to the row r of B. Step S2.2b: d^r(t_s^r) is corrected by the nonlinear correction module. Step S2.3: Then finally compute the sharp lines in t_s^r as L^r(t_s^r)≈B'^r / d^r(t_s^r).

[0088] These steps in this second exemplary method are repeated until the last row of B is processed.

[0089] In the above exemplary method, an intermediate resolution and color restoration submodule is used. Meanwhile, compared with the exemplary embodiment shown in Fig. 2, in order to use an event-based resolution and color correction submodule, in the third exemplary method shown in Fig. 4, the intermediate resolution and color restoration submodules in steps S2.2 and S2.5 are deleted from step S2, and instead, the event-based resolution and color correction submodule in step S1b is applied before step S2.1 of computing e^r^e(t_s^r). Specifically, as shown in Fig. 4, the event-based resolution and color correction submodule in step S1b is applied before step S2 to obtain E', i.e., e' described above in the "Event-Based Resolution and Color Restoration Submodule" section, and in step S2.1, E' is used to estimate the corresponding row r^e in the event stream, and row d^r^e(t_s^r) is computed from the event stream.

[0090] As another variation, FIG. 5 illustrates a fourth exemplary method according to the present invention, in which a post-resolution and color correction submodule is used. In this case, compared with the exemplary embodiment shown in FIG. 2, the intermediate resolution and color restoration submodules in steps S2.2 and S2.5 are deleted, image B is downscaled to the resolution of the event, the intermediate resolution and color restoration submodules are replaced with simple interpolation in steps S2.2′ and S2.5′, and step S2b is added to apply the post-resolution and color correction module. Specifically, as shown in FIG. 5, the post-resolution and color correction submodule step S2b is applied after step 2 to convert grayscale L to color L′.

[0091] In addition, the present invention also proposes several embodiments of frame enhancement using a fully data-driven approach. Specifically, the present invention proposes deblurring, creating multiple frames from a single blurred frame, or enhancing the dynamic range of a frame using a machine learning model in an end-to-end manner. Using the additional information provided by event-based sensors can reduce the model size and therefore the computational complexity.

[0092] That is, the present invention proposes to use machine learning to model: - Transformation of a blurred frame B and a stream of events (e_i)_i=1..N within or outside its exposure duration into a sharp image S at a given time within or outside the exposure duration of the blurred frame. - Transformation of a blurred frame B and the stream of events (e_i)_i=1..N during its exposure duration into a set of N sharp frames (S_0,…,S_N) regularly spaced in time throughout the exposure duration of the blurred frame. - transformation of a frame F1 with its exposure duration or any pre-processing of this stream of events (e_i)_i=1..N into a frame F2 with enhanced dynamic range.

[0093] The present invention also proposes a modification to the above mapping, where the event buffers (e_i)_i=1..N are replaced by any suitable pre-processing.

[0094] Additionally, the above-described modules, sub-modules and techniques according to the present invention are capable of obtaining slow motion video, enhanced dynamic range video in some further embodiments, as will be described below.

[0095] For slow-motion video, the exemplary method described above can generate an artifact-free image at a specific time t_ref. However, it is also possible to generate multiple corrected images at the output for each image at the input. Therefore, an embodiment according to the present invention can generate several frames at the output to increase the frame rate of the generated video, as shown in the following Figure 6.

[0096] Figure 6 shows the exposure window for two consecutive images captured by a frame-based camera. Similar to Figure 1, it consists of diagonal black elements representing the exposure window for each row of the rolling shutter, transformed into multiple compensated exposure windows, i.e., light gray vertical bars, representing 13 frames at 13 different times, thereby significantly increasing the frame rate.

[0097] Therefore, it is possible to generate high frame rate video without the constraints on exposure duration that occur when capturing high frame rate video with a frame camera alone (exposure duration must be less than 1 / (number of frames per second)). By using the present invention, it is possible to first capture, for example, a 60 frame / second video (so the maximum exposure is 16 ms) and then generate a video with a different (usually higher) number of frames / second, the maximum number of frames / second being constrained only by the temporal resolution of the event camera.

[0098] In the case of videos with enhanced dynamic range, in addition to the above-described embodiments aimed at enhancing the generation of a stream of frames and events that only contain other frames in the input, it is also possible to align and merge several images (or frames) captured sequentially by a sensor, and using several images corrected / enhanced by the present invention, it is possible to improve the quality of the images or videos and extend their dynamic range.

[0099] This improves upon methods in the art in which several frames are captured and spatially aligned before being averaging to reduce image noise. Alignment is difficult to obtain when the scene changes during frame capture due to moving objects in the scene or relative motion between the scene and the camera. Methods in the art generally use computationally intensive motion estimation and motion compensation to align frames. The present invention uses an event stream and image time shifting module to "align" several frames to a unique t_ref in a more efficient manner without explicit motion compensation.

[0100] The present invention is particularly advantageous when exposure durations vary between frame captures. Varying exposure durations is often used to better capture the dynamic range of a scene; when the highest and lowest brightnesses of a scene exceed the highest and lowest light intensities that can be represented in a single exposure due to the physical limitations of a (frame-based) sensor, it is common to capture N frames with different exposure durations E_i,i\in [0,N) and merge them into one "composite" frame, thus combining both the highest and lowest light intensities into a single frame. The varying exposure of the N frames makes alignment even more challenging for motion compensation techniques in the art.

[0101] In this case, the modules and methods described above can be used to capture several images of a scene and align them to a unique target time t_ref to obtain a high dynamic range (HDR) result.

[0102] Therefore, the present invention proposes using a combination of a conventional frame-based image sensor and an event-based sensor such as a DVS sensor to generate a stream of image / video frames and events, and then combining the image / video frames with the event data to more efficiently remove artifacts and enhance the dynamic range.

[0103] Moreover, the above-described modules / methods according to the present invention may be implemented in many ways, such as program instructions for execution by a processor, as software modules, microcode, as a computer program product on a computer-readable medium, as logic circuitry, as an application specific integrated circuit, as firmware, etc. Embodiments of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment containing both hardware and software elements. In a preferred embodiment, the present invention is implemented in software, which includes, but is not limited to, firmware, resident software, microcode, etc.

[0104] Furthermore, embodiments of the present invention may take the form of a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer, processing device, or any instruction execution system. For purposes of this description, a computer-usable or computer-readable medium may be any apparatus that can contain, store, communicate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The medium may be an electronic, magnetic, optical, or semiconductor system (or apparatus or device). Examples of computer-readable media include, but are not limited to, semiconductor or solid-state memory, magnetic tape, removable computer diskettes, RAM, read-only memory (ROM), rigid magnetic disks, optical disks, etc. Current examples of optical disks include compact disk-read-only memory (CD-ROM), compact disk-read / write (CD-R / W), and DVD.

[0105] The embodiments described hereinabove are illustrative of the present invention, and various modifications can be made thereto without departing from the scope of the invention which arises from the appended claims. [Explanation of symbols]

[0106] B Image frame

Claims

1. - capturing a reference image frame of the scene, the reference image frame including the artifact and having an exposure duration, by a frame-based rolling shutter sensor including rows of sensor elements; - capturing a stream of events by an event-based sensor synchronized with the frame-based rolling shutter sensor during at least the exposure duration, the events encoding brightness changes of the scene corresponding to the reference image frame; - processing each row of the reference image frame using the stream of events to derive the artifact-free corrected image frame; 1. A method for enhancing an image, comprising:

2. 2. The image enhancement method of claim 1, wherein the reference image frame is acquired at a reference time (t_r), and the step of deriving a corrected image frame includes a step of image time shifting, wherein the reference image frame is shifted to the corrected image frame at a target time (t) within or other than the exposure duration by establishing a relationship between the event and the reference image frame and accumulating the luminance change encoded by the event acquired during a time interval [t_r, t] between the reference time and the target time.

3. 2. The image enhancement method of claim 1, wherein the reference image frame is a blurred frame, and the method includes a step of image deblurring that transforms the blurred frame into a sharp frame.

4. The image enhancement method of claim 1 , further comprising the step of reblurring the corrected image frame.

5. The image enhancement method of claim 1 , further comprising a nonlinear compensation step for compensating for nonlinear processing in the reference image frame and the event.

6. 6. The image enhancement method of claim 1, wherein the reference image frame has a first resolution and / or a first color space and the stream of events has a second resolution and / or a second color space different from the first resolution and / or first color space, and the method further comprises a resolution and color restoration step to compensate for differences in resolution and / or color space between the reference image frame and the stream of events.

7. The image enhancement method of claim 1 , comprising generating a plurality of corrected image frames distributed at different times within or outside the exposure duration.

8. The image enhancement method of claim 1 , comprising registering a plurality of corrected image frames derived at different exposure durations.

9. - a frame-based rolling shutter sensor including rows of sensor elements configured to acquire a reference image frame having an exposure duration that includes the artifact; an event-based sensor synchronized with the frame-based rolling shutter sensor, configured to capture a stream of events encoding luminance changes of a scene corresponding to the reference image frame during at least the exposure duration; - a processor programmed to process each row of the reference image frame using the stream of events to derive the artifact-free corrected image frame; An image enhancement system comprising:

10. 10. The image enhancement system of claim 9, wherein the reference image frame is acquired at a reference time, and the processor is programmed to time-shift the reference image frame to the corrected image frame at the target time within or outside the exposure duration by establishing a relationship between the event and the reference image frame and accumulating the luminance change encoded by the event obtained during a time interval between the reference time and a target time.

11. 10. The image enhancement system of claim 9, wherein the processor is programmed to compensate for non-linear processing in the reference image frame and the event.

12. 10. The image enhancement system of claim 9, wherein the reference image frame has a first resolution and / or a first color space, and the stream of events has a second resolution and / or a second color space that is different from the first resolution and / or the first color space, and the processor is further programmed to compensate for differences in resolution and / or color space between the reference image frame and the stream of events.

Citation Information

Patent Citations

  • Real-time video deblurring

    EP2454876B1

  • Storing apparatus and storing method therefor, method of operating vision sensor, and computer-readable storage medium

    JP2017085549A

  • Imaging system and object recognition system

    JP2020161992A

  • Event camera

    JP2020182122A

  • Systems and methods for rolling shutter compensation using iterative process

    US10194101B1