Long-short exposure fusion with event data for low-light video enhancement

A neural network architecture combining long and short exposure RGB video streams with event data from event cameras addresses low-light video challenges by enhancing video quality through adaptive feature fusion and alignment, achieving improved denoising and deblurring.

WO2026161342A1PCT designated stage Publication Date: 2026-07-30DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2026-01-20
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing video capture technologies in low-light conditions face challenges with long exposure frames causing motion blur and low signal-to-noise ratio (SNR) issues, while short exposure frames suffer from noise and unreliable color due to image sensor limitations, and current techniques lack effective methods for enhancing video quality in such scenarios.

Method used

A neural network-based architecture combining long and short exposure RGB video streams with event data from event cameras, utilizing U-Net encoders and decoders, along with an SNR map and deformable convolutions, to adaptively fuse features and enhance video quality.

Benefits of technology

The proposed method effectively denoises and deblurs low-light video frames, improving framerate and SNR, resulting in sharper and more reliable video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2026011794_30072026_PF_FP_ABST
    Figure US2026011794_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods described herein provide low-light image and video enhancement. One example method for low-light image enhancement includes receiving a short-exposure input, a low-exposure input, and event data. The method includes brightening the short-exposure input based on a mean value of the long-exposure input to generate a brightened short-exposure image and calculating an SNR map based on the brightened short-exposure image. The method includes performing adaptive feature fusion based on the SNR map and the event data to generate a plurality of selected features and providing the plurality of selected features to a decoder.
Need to check novelty before this filing date? Find Prior Art

Description

LONG-SHORT EXPOSURE FUSION WITH EVENT DATA FOR LOW-LIGHT VIDEO ENHANCEMENT1. Cross-Reference to Related Applications

[0001] This application claims the benefit of U.S. Provisional Patent Application No.63 / 747,846, filed January 21, 2025, U.S. Provisional Patent Application No. 63 / 797,298, filed April 30, 2025, and European Patent Application No. EP 25176299.3, filed on May 14, 2025, each of which is incorporated by reference herein in its entirety.2. Field of the Disclosure

[0002] Various example embodiments relate to low-light video enhancement utilizing shortexposure inputs, long-exposure inputs, and event data.SUMMARY

[0003] In low-light situations, available options to taking a video sequence with consumer devices are limited. With the limitation of environmental brightness, users typically capture either a long exposure frame sequence or a short exposure frame sequence. For long exposure, where individual frame signal-to-noise ratio (SNR) is anticipated, there may be noticeable blurriness from both object movement and handshake of the user. The framerate is also limited to one over exposure time. For short exposure sequences, while framerate is guaranteed and there is no motion blur, the SNR of individual frames is low, leading to noisy frames with unreliable color due to limitations of image sensors. When capturing still images, prior techniques enhance camera performance by capturing two images, one for long exposure and another for short exposure, which are fused into a sharp low-noise image. However, few techniques exist for video.

[0004] Examples described herein provide systems and methods that combine long and short exposure RGB video streams with event data from event cameras to enhance the quality of video content. In some examples, a neural network-based architecture is provided that includes three encoders. The encoders include a short exposure encoder, a long exposure encoder, and an event voxels encoder. Decoders described herein may adaptively receive combined features skip-connected from the three encoders.

[0005] Examples described herein may also provide a mechanism for lightening the short exposure inputs using information from long exposure frames. In some instances, an SNR map is provided which is derived from the long exposure frames to guide the adaptive feature fusion for the decoder. A deformable convolution block may be provided in some examples to align the long exposure frames with multiple short exposure frames.

[0006] One example method for low-light image enhancement includes receiving a shortexposure input, a low-exposure input, and event data. The method includes brightening the shortexposure input based on a mean value of the long-exposure input to generate a brightened shortexposure image and calculating an SNR map based on the brightened short-exposure image. The method includes performing adaptive feature fusion based on the SNR map and the event data to generate a plurality of selected features and providing the plurality of selected features to a decoder.

[0007] Another example method for low-light image enhancement includes receiving, by a feature selection module, a first input associated with a short-exposure image, receiving, by the feature selection module, a second input associated with a long-exposure image, and receiving, by the feature selection module, a third input associated with event data. The method includes generating a feature based on the first input, the second input, and the third input, and providing the feature to a U-Net decoder configured to generate an output image.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which:

[0009] FIG. 1 illustrates an example of a long exposure video frame and a short exposure video frame taken in the same scene, according to some aspects.

[0010] FIG. 2 illustrates an example U-Net architecture, according to some aspects.

[0011] FIG. 3 illustrates an example architecture of Retinexf ormer, according to some aspects.

[0012] FIG. 4 illustrates an example architecture of EvLowLight, according to some aspects.

[0013] FIG. 5 illustrates an example architecture of REFID, according to some aspects.

[0014] FIG. 6 illustrates an example blur generation pipeline, according to some aspects.

[0015] FIG. 7 illustrates an example combination pipeline, according to some aspects.

[0016] FIG. 8 illustrates an example pipeline that combines long and short exposure RGB video streams with event data from event cameras, according to some aspects.

[0017] FIG. 9 illustrates example visual results of a blurry input, a noisy input, and an initial lighten-up result, according to some aspects.

[0018] FIG. 10 illustrates an initial brightened image and a corresponding SNR map, according to some aspects.

[0019] FIG. 11 illustrates an example Res block implemented in U-Net encoders, according to some aspects.

[0020] FIG. 12A illustrates an example of a standard convolution, according to some aspects.

[0021] FIG. 12B illustrates an example of a deformable convolution, according to some aspects.

[0022] FIG. 13 illustrates an example attention network, according to some aspects.

[0023] FIG. 14 illustrates a block diagram of a method for enhancing image and video frames, according to some aspects.

[0024] FIG. 15 illustrates a block diagram of another method for enhancing image and video frames, according to some aspects.

[0025] FIG. 16 illustrates another example pipeline based on an alignment module, according to some aspects.

[0026] FIG. 17 illustrates a block diagram of an example computing device, according to some aspects.

[0027] FIG. 18 shows various iterations for outputs of the retrained EvLowLight, according to some aspects.

[0028] FIGS. 19-30 illustrate example comparisons of the pipeline of FIG. 8 and baseline models for various images.DETAILED DESCRIPTION

[0029] In low-light situations, available options to taking a video sequence with consumer devices are limited. With the limitation of environmental brightness, users typically capture either a long exposure frame sequence or a short exposure frame sequence. For long exposure, where individual frame signal-to-noise ratio (SNR) is anticipated, there may be noticeable blurriness from both object movement and handshake of the user. The framerate is also limited to one-overexposure time I time exposure). For short exposure sequences, while framerate is guaranteed and there is no motion blur, the SNR of individual frames is low, leading to noisy frames with unreliable color due to limitations of image sensors.

[0030] FIG. 1 illustrates an example of a long exposure video frame 100 and a short exposure video frame 102 taken in the same scene. The long exposure video frame 100 has noticeable blurriness compared to the short exposure video frame 102. However, the short exposure video frame 102 includes more noise compared to the long exposure video frame 100.

[0031] When capturing still images, prior techniques enhance camera performance by capturing two images, one for long exposure and another for short exposure, which are fused into a sharp low-noise image. However, few techniques exist for video. Existing video techniques in low-light situations attempt to reduce motion blur of high-quality camera frames by using information from low-quality camera frames.

[0032] Event cameras are cameras that utilize CMOS sensors. Rather than capturing brightness intensity over a period of time per pixel location, event cameras take note of the timestamp when the brightness at the pixel location changes beyond a certain amount by the polarity. So, data from the event camera consists of a four-value array including time, polarity, and two-dimensional pixel location. The timestamp between two records may be so small that the framerate may be up to millions of frames per second, where each frame is an array representing a single pixel’s change.

[0033] Event cameras may be utilized as standalone cameras or combined with RGB cameras. Due to their high framerate and the ability to capture pixel intensity changes along with polarity, usage of event cameras for event-based motion estimation is increasing. One coherent event guidedlow-light video enhancement algorithm, referred to as “EvLowLighf utilizes event data to generate pixel-wise correlation between multiple low-light noisy frames in order to explore temporal correlations to build a background for frame denoising. Another algorithm, referred to as “REFID”, utilizes intensity information from event data to recurrently generate frames between two blurry frame inputs.

[0034] U-Net, an example architecture 200 of which is shown in FIG. 2, may be implemented in image restoration algorithms. The U-Net architecture 200 includes an encoding path 202 (or downsampling branch) and a decoding path 204 (or up-sampling branch) connected by a bottleneck 206. The encoding path 202 includes encoder layers that capture contextual information and reduce the spatial resolution of input frames at each layer. The decoding path 204 includes decoder layers that increases the spatial resolution of received frames at each layer while using the contextual information from the respective encoder layers. Skip connections between the encoding path 202 and the decoding path 204 indicated by the horizontal arrows are added at different scales of the latent features. In some examples, self-attention layers are inserted between convolution layers. However, known algorithms do not provide dual camera joint video denoising or deblurring.

[0035] Examples described herein provide systems and methods that combine long and short exposure RGB video streams with event data to enhance the quality of video content. In some examples, a neural network-based architecture is provided that includes three U-Net encoders. The encoders include a short exposure encoder, a long exposure encoder, and an event voxels encoder. Decoders described herein may be U-Net decoders that adaptively receive combined features skip-connected from the three encoders. While examples described herein primarily refer to U-Net architectures, other neural networks may also be utilized, such as convolutional neural networks (CNN), recurrent neural networks (RNN), or the like.

[0036] Examples described herein may also provide a mechanism for lightening the short exposure inputs using information from long exposure frames. In some instances, an SNR map is provided which is derived from the long exposure frames to guide the adaptive feature fusion for the decoder. A deformable convolution block may be provided in some examples to align the long exposure frames with multiple short exposure frames.

[0037] Examples described herein may be compared to different baseline architectures. For example, Retinexformer is a transformer-based low light enhancement network that does not use event data for denoising. FIG. 3 illustrates an example architecture 300 of Retinexformer.

[0038] As noted, event cameras may be implemented to provide motion information at high-temporal resolutions. EvLowLight utilizes event data to temporally and spatially align multiple input noisy frames by utilizing deformble convolutions with optical flow generated from event data (for example, by “EvFlow” networks). The alignment occurs at the lowest resolution of the image pyramid (e.g., the bottleneck connection). At the decoder, each level receives up-sampled features from the previous level and un-aligned features from encoders. The architecture 400 of EvLowLight is shown in FIG. 4. EvLowLight may be a U-Net based method with skip connection ignored. However, EvLowLight does not utilize exposures from event data and is less than desirable when the input images are noisy. EvLowLight may be implemented to denoise input images.

[0039] Known video deblurring methods improve blurriness of the long exposure frames, but do not compensate for low framerates. REFID is an example one-stage video deblurring and frame interpolation architecture. REFID is a recurrent U-Net based architecture where blurry frames and events pass through a one-pass encoder, while event voxels pass through another recurrent-based encoder based on how many frames are to be interpolated. The background may be extracted from the one-pass encoder and per-frame cumulative changes may be extracted from the recurrent encoder. The architecture 500 of REFID is shown in FIG. 5. REFIED may be implemented to deblur input images.

[0040] Examples described herein may utilize data from three cameras: a first RGB camera providing high-framerate, short-exposure noisy frames, a second RG camera providing low-framerate, long-exposure blurry frames, and one event camera providing event data. In other implementations, fewer or additional cameras may be utilized, such as additional event cameras. Denoising and deblurring separate to fuse intermediate outcomes relies on the performance of separate denoising / deblurring networks and may be inconsistent temporally. Examples described herein provide a pipeline that includes initial-relighting, SNR map estimation, feature extraction, feature alignment, and feature fusion for low-light long- short-event fusion. The input data set may include approximately 91 video clips of around 300 frames each clip, including indoor and outdoorscenes. RGB images and events may have a resolution of approximately 346x260 in examples described herein. However, other resolutions are contemplated. The three data streams may be aligned temporally.

[0041] To handle blurry frames, examples described herein interpolate original video (for example, with three additional frames in between, four frames in between, or the like), then averages the video every set of frames (for example, every six frames, seven frames, eight frames, or the like). Additionally, to reduce over-exposure from averaging frames, a saturation mask may be applied to randomly add over-exposure based on a fraction of original frames over-exposure at each pixel location. An example of the blur generation pipeline 600 is shown in FIG. 6. The blur generation pipeline 600 may be implemented to generate training images for training neural network architectures described herein.RGB Fusion Methods

[0042] In some implementations, the outcomes from deblurring and denoising networks are fused. For example, the deblurring and denoising outcomes may be weighted averaged by an attention map calculated by a post-trained U-Net, which receives the original images, the outcome images, and the events as inputs.

[0043] FIG. 7 illustrates an example combination pipeline 700. The combination pipeline 700 receives short-exposure input frames 702, long-exposure input frames 704, and event camera data 706. The short-exposure input frames 702 are input to a denoising module 708 configured to denoise the short-exposure input frames 702. In some examples, the denoising module 708 is the EvLowLight architecture 400 of FIG. 4 The long-exposure input frames 704 are input to a deblurring module 710 configured to deblur the long-exposure input frames 704. In some examples, the denoising module 710 is the REFID architecture 500 of FIG. 5.

[0044] The event camera data 706 is input to a fusion U-Net architecture 714 as event voxels 712. The event voxels 712 may be four-value array including time, polarity, and two-dimensional pixel location. The U-Net architecture 714 calculates an attention map 716 that performs a weighted average of an output of the denoising module 708 and the deblurring module 710. The weighted average performed by the attention map 716 results in output frames 718 that are output by the combination pipeline 700.

[0045] In other examples described herein, an end-to-end approach is provided. For example, both long and short exposure images are implemented in the encoder path to extract global information and maintain low-frequency information of the scene. High-frequency information is added to the image in the decoders from a combination of skip connection from the blur and noise encoders, and a separate voxel feature extractor, guided by the SNR map calculated from the noisy image.

[0046] FIG. 8 illustrates an example pipeline 800 that combines long and short exposure RGB video streams with event data from event cameras to enhance the quality of video content. For example, the pipeline 800 includes long-exposure input frames 802, short-exposure input frames 804, and event voxels 806. In some examples, the long-exposure input frames 802 are received from a first camera, the short-exposure input frames 804 are received from a second camera, and the event voxels 806 are received from an event camera. The event voxels may be a four-value array including time, polarity, and two-dimensional pixel location, and may indicate changes in pixel value. In other examples, the long-exposure input frames 802 and the short-exposure input frames 804 are received from a same camera. In yet another example, the long-exposure input frames 802, the short-exposure input frames 804, and the event voxels 806 are received from a database.

[0047] In some instances, the event voxels are represented as a grid and a 3D histogram of individual events (corresponding to both position and time) is created. In such an instance, each event’s polarity is spread among its closest voxels using a kernel and may be represented according to Equation (1):V(x,y,t) = iPt^bCx - Xi)kb(y - yi)kb(t - t-) Equation (1)Where:kb(a) = max(0, 1 — |a|);where B is the number of bins implemented to discretize the time dimension when the event sequence has M events within the time interval of [tq, tM],

[0048] The pipeline 800 includes a brightness scale module 808. Since the blurry image is a normalized temporal integration of the ground truth, the blurry image maintains the most accurateglobal color and illumination information among the three input sources. While prior techniques such as EvLow Light and EvLight estimate exposure information from dark frames, examples described herein instead get exposure information directly from the blurry frame. For example, the initial-light-up image may be represented according to Equation (2):Equation (2)where ILUrepresents the outcome of the lighten-up image, Tbrepresents the mean value of the long-exposure (blurry) frame, andnrepresents the mean value of the short-exposure (noisy) frame. For an example input batch (one long exposure frame and N short exposure frames), each short exposureframe is lightened-up by the brightness scale module 808 with independent scale where i ~ [O^V).NThe brightness scale module 808 outputs brightened short exposure frames 810.

[0049] FIG. 9 illustrates example visual results of the blurry input 900, the noisy input 902, and the initial lighten-up result 904 (which may be, for example, the brightened short exposure frame 810). In some instances, the initial lighten-up may be performed by a neural network.

[0050] The pipeline 800 includes an SNR map estimation module 812. The SNR map estimation module 812 receives the brightened short exposure frames 810 from the brightness scale module 808. The SNR map estimation module 812 calculates an SNR map 814 based on the initial brightened short exposure frame 810 (JLU) in order to guide the feature selection module for image reconstruction at the decoder. For example, an initial denoised image Idis acquired by applying a low-pass filter (e.g., a Gaussian filter or a 5*5 box filter) to the initial brightened short exposure frame 810. The SNR is calculated by the SNR map estimation module 812 by dividing the signal with the noise. The SNR is mathematically represented according to Equation (3):where ® represents convolution and k represents a blur kernel, for example a low pass blur kernel. FIG. 10 illustrates visual results of an initial brightened image 1000 and a corresponding SNR map 1002. In another example, the SNR is calculated according to Equation (4):Equation (4)where U is the initial brightened short exposure frame 810 and U’ is the initial brightened short exposure frame 810 convolved with the filter k.

[0051] As previously mentioned, U-Net is an encoder-decoder structure with skip-connection that passes high-frequency features from encoder to decoder. The pipeline 800 includes three encoders: a long encoder 816. a short encoder 818, and an event encoder 820. In the encoder phase of the pipeline 800, each encoder includes multiple levels. The long encoder 816, the short encoder 818, and / or the event encoder 820 may be U-Net encoders (e.g., the encoding path 202) that are connected to the decoder 826 (e.g., the decoding path 206) via the feature selection module 824. At each level, the resolution of the feature is downsampled by pooling or stride convolution. While the structure of each encoder is similar, the blur path concatenates features from the noise path before the convolution layer to extract holistic features and build up solid background features. The outputs of the short encoder 818 and the event encoder 820 are provided to a feature selection module 824. The output of the long encoder 816 is provided to a DCN alignment module 822, which aligns the features of the output of the long encoder 816 before passing the aligned features to the feature selection module 824. The feature selection module 824 buffers the outputs of the encoders and provides an input to a decoder 826.

[0052] The long encoder 816 receives the long-exposure input frames 802 as inputs. The short encoder 818 receives the brightened short exposure frames 810 as inputs. The event encoder 820 receives event voxels 806 as inputs. For example, events detected by an event camera may be converted to the event voxels 806 for processing.

[0053] FIG. 11 illustrates an example Res block 1100 implemented at each level of the long encoder 816, the short encoder 818, and / or the event encoder 820. The Res block 1100 includes a first module 1102 and a second module 1104. The first module 1102 is a sequential of a 2D convolution layer, a batch normalization layer, a Relu activation function, another 2D convolution layer, and another batch normalization. The first module 1102 computes the residual to be added back to the input features. The result of addition is provided to the second module 1104, which includes a 2D convolution layer, a batch normalization layer, and a Relu activation function. The resulting feature is passed to the feature selection module 824 (shown in FIG. 8). where it is buffered to be combined and passed to the decoder 826. For each encoder, there may be four levelsof the feature pyramid, with each level receiving a max pooled feature of the previous level (and the feature resolution being half the previous level).

[0054] The pipeline 800 also includes a DCN alignment module 822, or a blur feature alignment module. With three primary input sources of exposure (long exposure, short exposure, and event), one long exposure frame may correspond to multiple noisy frames and multiple event data.Accordingly, the decoder may (i) receive an aligned version of the blur features at each output frame location and / or (ii) use blur features only when necessary to minimize the misaligned feature resulting in artifacts. When receiving the aligned version of the blur features, blur feature alignment may be implemented by the DCN alignment module 822. When blur features are only used when necessary, feature selective fusion may be implemented.

[0055] For feature alignment performed by the DCN alignment module 822, a deformable convolution may be utilized for the skip-connected blur features. FIG. 12A illustrates an example of a standard convolution. FIG. 12B illustrates an example of the deformable convolution. The input of the DCN alignment module 822 may correspond to the noisy features, blurry features, and offset calculated from the previous level. The offset of the lowest level of the pyramid may be none, while the upper level may receive upscaled offsets from lower levels.

[0056] For feature selection performed by the feature selection module 824, the SNR map 814 is implemented to guide feature selection for an improved decoding result. Feature selection may include a pixel- wise multiplication or an attention network. FIG. 13 illustrates an example attention network 1300. In the example of FIG. 13, the features input to the attention network 1300 are the outputs of the DCN alignment module 822, the short encoder 818, and the event encoder 820. The input SNR map 814 may be first passed through a tanh function to stretch the SNR map 814 to a range of (0, 1) before multiplying the skip-connected noise feature with the SNR map 814 and the blur and event feature with (1-SNR map).

[0057] In some examples, there is one path for the decoder, where in each level the input is from the SNR-based feature selection module 824 and the upscale feature from the previous level is concatenated. Feature upscaling may be performed by an upsampling module followed with a 2D convolution. The concatenated features may pass through a Res block 1100, previously shown in FIG. 11. The output of the Res block 1100 may be upscaled in the next level in the decoder 826. The decoder 826 outputs an output frame 828. In the final stage, the output frame 828 may bepassed to a 1x1 convolution layer that converts the final layer channel to three channels as RGB outputs.

[0058] FIG. 14 illustrates a block diagram of an example method 1400 for enhancing image and video frames. The method 1400 may be performed by, for example, the pipeline 800. The steps provided within FIG. 14 are merely examples, and may instead be conducted in a different order or simultaneously. In some examples, steps may be omitted, or additional steps may be provided.

[0059] At block 1402, the pipeline 800 receives a short-exposure input, a long-exposure input, and event data. For example, the pipeline 800 receives the long-exposure input frames 802, the short-exposure input frames 804, and event voxels 806.

[0060] At block 1404, the pipeline 800 brightens the short-exposure input based on a mean value of the long-exposure input to generate a brightened short-exposure image. For example, the brightness scale module 808 brightens the short-exposure input frames 804 using independent scale to generate brightened short exposure frames 810, as previously described with respect to Equation (2).

[0061] At block 1406, the pipeline 800 calculates an SNR map based on the brightened shortexposure image. For example, the SNR map estimation module 812 calculates an SNR map 814 based on the initial brightened short exposure frame 810 (JLU), as previously described with respect to Equation (3).

[0062] At block 1408, the pipeline 800 performs adaptive feature fusion based on the SNR map to generate a plurality of selected features. For example, the feature selection module 824 receives the SNR map 814 alongside outputs from the long encoder 816, the short encoder 818, and the event encoder 820. The feature selection module 824 may implement an attention network 1300 to generate a plurality of selected features.

[0063] At block 1410, the pipeline 800 provides the plurality of selected features to a decoder. For example, the feature selection module 824 provides the plurality of selected features to the decoder 826.

[0064] FIG. 15 illustrates a block diagram of an example method 1500 for enhancing image and video frames. The method 1500 may be performed by, for example, the pipeline 800 (and, moreparticularly, the feature selection module 824). The steps provided within FIG. 15 are merely examples, and may instead be conducted in a different order or simultaneously. In some examples, steps may be omitted, or additional steps may be provided. For example, the method 1500 may be performed in conjunction with the method 1400 of FIG. 14.

[0065] At block 1502, the feature selection module 824 receives a first input associated with a short-exposure image. For example, the feature selection module 824 receives an output from the short encoder 818 associated with the brightened short exposure frames 810.

[0066] At block 1504, the feature selection module 824 receives a second input associated with a long-exposure image. For example, the SNR map 814 receives an output from the long encoder 816 associated with the long-exposure input frames 802.

[0067] At block 1506, the feature selection module 824 receives a third input associated with event data. For example, the feature selection module 824 receives an output from the event encoder 820 associated with the event voxels 806.

[0068] At block 1508, the feature selection module 824 generates a feature based on the first input, the second input, and the third input. For example, the feature selection module 824 implements the attention network 1300 to generate a feature.

[0069] At block 1510, the feature selection module 824 provides the feature to a U-Net decoder configured to generate an output image. For example, the feature selection module 824 provides the feature to the decoder 826, which generates output frame 828.

[0070] FIG. 16 illustrates another example pipeline 1600 based on an alignment module (propagation). In the pipeline 1600 of FIG. 16, blur features are also passed to the decoder 826 using a similar alignment module, but flow is reversed compared to the pipeline of FIG. 8, going from a long exposure frame to a different short exposure frame location. The pipeline 1600 of FIG.14 may not include exposure information from event voxels 806.

[0071] In the pipeline 1600, the DCN alignment module 822 is replaced with an EvFlow module 1602 and a flow-based propagation module 1604. The EvFlow module 1602 estimates the optical flow based on event data, using the event voxels 806 as an input. The flow-based propagation module 1604 may be, for example, EvLowLight, as previously described with respect to FIG. 4.The flow-based propagation module 1604 uses the optical flow computed by the EvFlow module 1602 and warps the features of the frames and the events. The flow-based propagation module 1604 predicts offsets to the optical flow based on these warped features. The output of the flow-based propagation module 1604 is provided to the feature selection module 824.

[0072] FIG. 17 is a block diagram of an example computing device 1700 one or more instances of which can be used to implement various above-described methods and workflows according to various examples.

[0073] The computing device 1700 of FIG. 17 is illustrated as having a number of components, but any one or more of these components may be omitted or duplicated, as suitable for the application and setting. In some embodiments, some or all of the components included in the computing device 1700 may be attached to one or more motherboards and enclosed in a housing. In some embodiments, some of those components may be fabricated onto a single system-on-a-chip (SoC) (e.g., the SoC may include one or more electronic processing devices 1702 and one or more storage devices 1704). Additionally, in various embodiments, the computing device 1700 may not include one or more of the components illustrated in FIG. 17, but may include interface circuitry for coupling to the one or more components using any suitable interface (e.g., a Universal Serial Bus (USB) interface, a High-Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, a wireless interface, or any other appropriate interface). For example, the computing device 1700 may not include a display device 1710, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which an external display device 1710 may be coupled.

[0074] The computing device 1700 includes a processing device 1702 (e.g., one or more processing devices). As used herein, the terms “electronic processor device” and “processing device” interchangeably refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that may be stored in registers and / or memory. In various embodiments, the processing device 1702 may include one or more digital signal processors (DSPs), application-specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), server processors, or any other suitable processing devices.

[0075] The computing device 1700 also includes a storage device 1704 (e.g.. one or more storage devices). In various embodiments, the storage device 1704 may include one or more memory devices, such as random-access memory (RAM) devices (e.g., static RAM (SRAM) devices, magnetic RAM (MRAM) devices, dynamic RAM (DRAM) devices, resistive RAM (RRAM) devices, or conductive-bridging RAM (CBRAM) devices), hard drive-based memory devices, solid-state memory devices, networked drives, cloud drives, or any combination of memory devices. In some embodiments, the storage device 1704 may include memory that shares a die with the processing device 1702. In such an embodiment, the memory may be used as cache memory and include embedded dynamic random-access memory (eDRAM) or spin transfer torque magnetic random-access memory (STT-MRAM), for example. In some embodiments, the storage device 1704 may include non-transitory computer readable media having instructions thereon that, when executed by one or more processing devices (e.g.. the processing device 1702). cause the computing device 1700 to perform any appropriate ones of the methods disclosed herein or portions of such methods.

[0076] The computing device 1700 further includes an interface device 1706 (e.g., one or more interface devices 1706). In various embodiments, the interface device 1706 may include one or more communication chips, connectors, and / or other hardware and software to govern communications between the computing device 1700 and other computing devices. For example, the interface device 1706 may include circuitry for managing wireless communications for the transfer of data to and from the computing device 1700. The term “wireless” and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data via modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not. Circuitry included in the interface device 1706 for managing wireless communications may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.11 family). IEEE 802.16 standards, Long-Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as “3GPP2”), etc.). In some embodiments, circuitry included in the interface device 1706 for managing wireless communications may operate in accordance with a Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access(HSPA), Evolved HSPA (E-HSPA), or LTE network. In some embodiments, circuitry included in the interface device 1706 for managing wireless communications may operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). In some embodiments, circuitry included in the interface device 1706 for managing wireless communications may operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G. and beyond. In some embodiments, the interface device 1706 may include one or more antennas (e.g., one or more antenna arrays) configured to receive and / or transmit wireless signals.

[0077] In some embodiments, the interface device 1706 may include circuitry for managing wired communications, such as electrical, optical, or any other suitable communication protocols. For example, the interface device 1706 may include circuitry to support communications in accordance with Ethernet technologies. In some embodiments, the interface device 1706 may support both wireless and wired communication, and / or may support multiple wired communication protocols and / or multiple wireless communication protocols. For example, a first set of circuitry of the interface device 1706 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second set of circuitry of the interface device 1706 may be dedicated to longer-range wireless communications such as global positioning system (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some other embodiments, a first set of circuitry of the interface device 1706 may be dedicated to wireless communications, and a second set of circuitry of the interface device 1706 may be dedicated to wired communications.

[0078] The computing device 1700 also includes battery / power circuitry 1708. In various embodiments, the battery / power circuitry 1708 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 1700 to an energy source separate from the computing device 1700 (e.g., to AC line power).

[0079] The computing device 1700 also includes a display device 1710 (e.g.. one or multiple individual display devices). In various embodiments, the display device 1710 may include anyvisual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.

[0080] The computing device 1700 also includes additional input / output (I / O) devices 1712. In various embodiments, the I / O devices 1712 may include one or more data / signal transfer interfaces, audio I / O devices (e.g., microphones or microphone arrays, speakers, headsets, earbuds, alarms, etc.), audio codecs, video codecs, printers, sensors (e.g., thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration sensors, etc.), image capture devices (e.g., one or more cameras), human interface devices (e.g., keyboards, cursor control devices, such as a mouse, a stylus, a trackball, or a touchpad), etc.

[0081] Depending on the specific embodiment, various components of the interface devices 1706 and / or I / O devices 1712 can be configured to output suitable control signals, receive suitable control / telemetry signals, and receive and transmit data streams. In some examples, the interface devices 1706 and / or I / O devices 1712 include one or more analog-to-digital converters (ADCs) for transforming received analog signals into a digital form suitable for operations performed by the processing device 1702 and / or the storage device 1704. In some additional examples, the interface devices 1706 and / or I / O devices 1712 include one or more digital-to-analog converters (DACs) for transforming digital signals provided by the processing device 1702 and / or the storage device 1704 into an analog form suitable for being transmitted through a communication channel.Implementation

[0082] EvLowLight, REFID, and Retinexformer are referenced as baseline algorithms for comparison. A modified SDE dataset may be implemented for training and validation of the baseline algorithms and examples described herein (e.g., the pipeline 800 of FIG. 8). The training set includes 75 sequences (35 indoors and 40 outdoors) and the testing set includes 16 sequences (8 indoors and 8 outdoors). Random crop, flip, and rotation are applied for data augmentation. There are seven ground truth frames, one long exposure frame, and seven short exposure frames and corresponding events for each iteration.

[0083] Additionally, EvLowLight was retrained on the modified SDE dataset with 280,000 iterations. The learning rate was set to 10’4and decay every 50,000 iterations. FIG. 18 shows acomparison from the 60.000thiteration (shown by first image 1800) to the 280.000thiteration (shown by second image 1805) for the retrained EvLowLight.

[0084] For training of the pipeline 800, a batch size of two with input size of 256*256 is implemented. For the validation set, the input image size is the same as the original image size. The same learning rate settings as EvLowLight were implemented, for 200.000 iterations, and continued for another 200,000 iterations if convergence was not satisfied. For the loss function, the average L2 loss between the outcome and the ground truth images per iteration was used.

[0085] FIGS. 19-30 illustrate example comparisons of the pipeline 800 and baseline models for various images. Table 1 provides numerical results (PSNR and SSIM) of the reconstruction for various models.Table 1: Numerical results in reconstruction

[0086] The present disclosure likewise relates to corresponding computer programs, computer program products, and computer-readable storage media storing such computer programs or computer program products. Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.

[0087] Aspects of the methods and apparatus / systems described herein may be implemented in an appropriate computer-based audio processing network environment (e.g., server or cloud environment) for processing digital or digitized audio files. Portions of the audio system may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among thecomputers. Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.

[0088] One or more of the components, blocks, processes or other functional components (modules) may be implemented through a computer program that controls execution of a processorbased computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and / or as data and / or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, physical (non-transitory), non-volatile storage media in various forms, such as optical, magnetic or semiconductor storage media.

[0089] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art. and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and / or application specific integrated circuits (“ASICs”). As such, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, the apparatus (e.g., encoders) described above can include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.

[0090] A person skilled in the art realizes that the present invention by no means is limited to the embodiments described above. On the contrary, many modifications and variations are possible and considered within the scope of the appended claims. Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims, and which may represent systems, methods, and devices, all arranged in accordance with aspects of the present disclosure.

[0091] EEE1. A method for low-light image enhancement, the method comprising: receiving a short-exposure input, a long-exposure input, and event data; brightening the short-exposure input based on a mean value of the long-exposure input to generate a brightened short-exposure image; calculating an SNR map based on the brightened short-exposure image; performing adaptive feature fusion based on the SNR map and the event data to generate a plurality of selected features; and providing the plurality of selected features to a decoder.

[0092] EEE2. The method of EEE1, wherein brightening the short-exposure input includes: calculating a scale by dividing the mean value of the long-exposure input by a mean value of the short-exposure input; and multiplying the short-exposure input by the scale.

[0093] EEE3. The method of any one of EEE1 to EEE2, wherein calculating the SNR map includes: applying a low-pass filter to the brightened short-exposure image to generate an initial denoised image; and dividing the initial denoised image with noise.

[0094] EEE4. The method of any one of EEE1 to EEE3, wherein the event data indicates pixel intensity changes along with polarity.

[0095] EEE5. The method of any one of EEE1 to EEE4, further comprising: aligning the shortexposure input and the long-exposure input using a deformable convolution.

[0096] EEE6. The method of any one of EEE1 to EEE5, wherein receiving the short-exposure input, the long-exposure input, and the event data includes receiving the event data from an event camera.

[0097] EEE7. The method of any one of EEE1 to EEE6, further comprising: providing the long-exposure input to a first U-Net encoding stage; providing the brightened short-exposure image to a second U-Net encoding stage: and providing the event data to a third U-Net encoding stage.

[0098] EEE8. The method of EEE7, wherein performing adaptive feature fusion based on the SNR map to generate the plurality of selected features includes: providing an output of the first U-Net encoding stage, the second U-Net encoding stage, and the third U-Net encoding stage to an adaptive feature fusion module configured to perform the adaptive feature fusion.

[0099] EEE9. The method of EEE8, wherein the adaptive feature fusion module includes an attention network.

[0100] EEE10. The method of any one of EEE1 to EEE9, wherein the decoder is a U-Net decoding stage.

[0101] EEE11. A method for low-light image enhancement, the method comprising: receiving, by a feature selection module, a first input associated with a short-exposure image; receiving, by the feature selection module, a second input associated with a long-exposure image; receiving, by the feature selection module, a third input associated with event data; generating a feature based on the first input, the second input, and the third input; and providing the feature to a U-Net decoder configured to generate an output image.

[0102] EEE12. The method of EEE11, further comprising: receiving the short-exposure image; brightening the short-exposure image to generate a brightened short-exposure image; and providing the brightened short-exposure image to a U-Net encoder configured to generate the first input.

[0103] EEE13. The method of any one of EEE11 to EEE12, further comprising: receiving the long-exposure image; and providing the long-exposure image to a U-Net encoder configured to generate the second input.

[0104] EEE 14. The method of any one of EEE 11 to EEE 13, further comprising: receiving the event data; and providing the event data to a U-Net encoder configured to generate the third input.

[0105] EEE15. The method of any one of EEE11 to EEE14, further comprising: receiving, by the feature selection module, an SNR map associated with the short-exposure image, wherein the feature is generated further based on the SNR map.

[0106] EEE16. The method of EEE15, further comprising: calculating the SNR map by applying a low-pass filter to a brightened short-exposure image, thereby generating an initial denoised image; and dividing the initial denoised image with noise.

[0107] EEE17. The method of any one of EEE11 to EEE16, further comprising: aligning the first input and the second input using a deformable convolution.

[0108] EEE18. The method of any one of EEE11 to EEE17, wherein generating the feature includes applying the first input, the second input, and the third input to an attention network.

[0109] EEE19. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEE1 to EEE18.

[0110] EEE20. A non-transitory computer-readable storage medium storing the program according to EEE 19.

[0111] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.

[0112] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.

[0113] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.

[0114] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used tointerpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.

[0115] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims.

[0116] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit.

[0117] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium including being loaded into and / or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits.

[0118] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word “about” or “approximately” preceded the value or range.

[0119] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.

[0120] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence.

[0121] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments necessarily mutually exclusive of other embodiments. The same applies to the term “implementation.”

[0122] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objects so referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner.

[0123] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if’ may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].”

[0124] Also, for purposes of this description, the terms “couple.” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developed in which energy is allowed to be transferred between two or more elements, and the interposition ofone or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements.

[0125] As used herein in reference to an element and a standard, the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard and would be recognized by other elements as sufficiently capable of communicating with the other elements in the manner specified by the standard. The compatible element does not need to operate internally in a manner specified by the standard.

[0126] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and / or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and / or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.

[0127] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) foroperation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

[0128] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.

Claims

CLAIMSWhat is claimed is:

1. A method for low-light image enhancement, the method comprising:receiving a short-exposure input, a long-exposure input, and event data;brightening the short-exposure input based on a mean value of the long-exposure input to generate a brightened short-exposure image;calculating an SNR map based on the brightened short-exposure image;performing adaptive feature fusion based on the SNR map and the event data to generate a plurality of selected features; andproviding the plurality of selected features to a decoder.

2. The method of claim 1, wherein brightening the short-exposure input includes:calculating a scale by dividing the mean value of the long-exposure input by a mean value of the short-exposure input; andmultiplying the short-exposure input by the scale.

3. The method of claim 1 or claim 2, wherein calculating the SNR map includes:applying a low-pass filter to the brightened short-exposure image to generate an initial denoised image; anddividing the initial denoised image with noise.

4. The method of any one of claims 1-3, wherein the event data indicates pixel intensity changes along with polarity.

5. The method of any one of claims 1-4, further comprising:aligning the short-exposure input and the long-exposure input using a deformable convolution.

6. The method of any one of claims 1-5, wherein receiving the short-exposure input, the long-exposure input, and the event data includes receiving the event data from an event camera.

7. The method of any one of claims 1-6, further comprising:providing the long-exposure input to a first U-Net encoding stage;providing the brightened short-exposure image to a second U-Net encoding stage; and providing the event data to a third U-Net encoding stage.

8. The method of claim 7, wherein performing adaptive feature fusion based on the SNR map to generate the plurality of selected features includes:providing an output of the first U-Net encoding stage, the second U-Net encoding stage, and the third U-Net encoding stage to an adaptive feature fusion module configured to perform the adaptive feature fusion.

9. The method of claim 8, wherein the adaptive feature fusion module includes an attention network.

10. The method of any one of claims 1-9, wherein the decoder is a U Net decoding stage.

11. A method for low-light image enhancement, the method comprising:receiving, by a feature selection module, a first input associated with a short-exposure image; receiving, by the feature selection module, a second input associated with a long-exposure image;receiving, by the feature selection module, a third input associated with event data; generating a feature based on the first input, the second input, and the third input; and providing the feature to a U-Net decoder configured to generate an output image, the method further comprising:receiving the short-exposure image;brightening the short-exposure image to generate a brightened short-exposure image; and providing the brightened short-exposure image to a U-Net encoder configured to generate the first input.

12. The method of claim 11, further comprising:receiving the long-exposure image: andproviding the long-exposure image to a U-Net encoder configured to generate the second input.

13. The method of claim 11 or claim 12, further comprising:receiving the event data; andproviding the event data to a U-Net encoder configured to generate the third input.

14. The method of any one of claims 11-13, further comprising:receiving, by the feature selection module, an SNR map associated with the short-exposure image,wherein the feature is generated further based on the SNR map.

15. The method of claim 14, further comprising:calculating the SNR map by applying a low-pass filter to a brightened short-exposure image, thereby generating an initial denoised image; anddividing the initial denoised image with noise.

16. The method of any one of claims 11-15, further comprising:aligning the first input and the second input using a deformable convolution.

17. The method of any one of claims 11-16, wherein generating the feature includes applying the first input, the second input, and the third input to an attention network.

18. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 1-17.

19. A non-transitory computer-readable storage medium storing the program according to claim 18.