Low-light image enhancement method based on event and frame

By employing a dual attention mechanism based on events and frames, efficient enhancement of low-light images is achieved, solving the problem of image quality degradation caused by noise fusion and generating clear images with high fidelity and high similarity.

CN121921222APending Publication Date: 2026-04-24BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING INST OF TECH
Filing Date
2025-11-21
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing low-light enhancement methods based on frames and events are ineffective in noise fusion, making it difficult to recover clear edge structures and details in low-light image reconstruction. Furthermore, traditional methods amplify noise during the enhancement process, affecting image quality.

Method used

A dual attention mechanism based on events and frames is adopted. By synchronously collecting event data and frame data, multi-scale feature encoding and dynamic weighted feature fusion are performed. The convolutional attention algorithm is used to adjust the modality weights and generate clear image reconstruction results.

Benefits of technology

It effectively suppresses noise, restores image details, improves image fidelity and structural similarity, significantly improves visual effects in low-light environments, and generates natural and clear images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921222A_ABST
    Figure CN121921222A_ABST
Patent Text Reader

Abstract

A low-light image enhancement method based on events and frames comprises the following steps: synchronously acquiring event data and low-light frame data of an event camera, and performing event stream preprocessing on the event data to convert the event data into five-channel event voxels; multi-scale feature coding based on a double-attention mechanism is carried out on the low-light frame data and the five-channel event voxels, image features are enhanced from channel dimensions and spatial dimensions, and enhanced feature maps corresponding to the low-light frame data and the event voxels are obtained; performing dynamic weighted feature fusion on the enhanced feature maps corresponding to the low-light frame data and the event voxels to obtain a fusion feature map of the event voxel information and the low-light frame data information; and performing up-sampling image reconstruction and enhancement on the fused feature map to obtain a final clear image. According to the method, multi-scale events and image features can be adaptively weighted and fused, noise is effectively suppressed, natural high-quality images with clear details are recovered, and the perception capability of a visual system in a low-light environment is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision and image processing technology, and in particular to an event- and frame-based low-light image enhancement method, which is a technique for image enhancement in low-light environments. Background Technology

[0002] High-resolution images provide clear detail and are fundamental to computer vision tasks. However, images captured in low-light environments often suffer from degraded image quality, such as low contrast, low visibility, strong noise, and inaccurate color. Reconstructing high-quality images from low-light input has become a key technical challenge in the field of computer vision.

[0003] Currently, there are several frame-based methods designed to improve the quality of images captured under low-light conditions. While these algorithms can improve overall brightness, the low signal-to-noise ratio of the original image often amplifies noise in dark areas during the enhancement process, making high-quality image reconstruction significantly difficult, especially when visual details such as edges captured by the frame camera are unclear.

[0004] Unlike traditional cameras that capture global images at a fixed frame rate, event cameras, as a novel type of biomimetic vision sensor, use pixels to sense changes in light intensity, asynchronously trigger events, and output an event stream to record scene information. This working method enables event cameras to have high dynamic range performance (up to 140dB, compared to 60dB for traditional cameras), allowing them to accurately capture the edge structure of dynamic scenes in low-light environments.

[0005] Leveraging the edge delineation advantages of event-based cameras, some event-based low-light enhancement studies have been proposed. However, event-based methods alone cannot obtain the initial pixel values ​​of an image, making it difficult to recover the absolute intensity of the scene. Considering the irreplaceable advantages of both modalities, event- and frame-based low-light enhancement studies have gradually emerged in recent years. However, since both modalities may exhibit different noise patterns, this could enhance noise fusion, thereby weakening effective feature representation. Summary of the Invention

[0006] Based on the above analysis, the present invention aims to provide a low-light image enhancement method based on event and frame dual attention mechanism, which effectively solves the problem of poor low-light enhancement effect caused by noise amplification and detail loss in the direct fusion method, and achieves excellent low-light image enhancement effect.

[0007] On one hand, embodiments of the present invention provide a low-light image enhancement method based on a dual attention mechanism of events and frames, the method comprising the following steps:

[0008] Simultaneously acquire event data generated by the event sensor in the event camera and low-light frame data generated by the APS frame sensor, and perform event stream preprocessing on the event data to convert the event data into five-channel event voxels;

[0009] Multi-scale feature encoding based on a dual attention mechanism is performed on the low-light frame data and the five-channel event voxels to enhance image features from both the channel dimension and the spatial dimension, thereby obtaining enhanced feature maps corresponding to the low-light frame data and event voxels.

[0010] Dynamically weighted feature fusion is performed on the enhanced feature maps corresponding to low-light frame data and event voxels. A convolution-based spatial-channel attention algorithm is used to adjust the weights of different modalities to obtain a fused feature map of the edge information of event voxels and the color and texture information of low-light frame data.

[0011] Image reconstruction and enhancement are performed by upsampling the fused feature map to obtain the final clear image.

[0012] Furthermore, the multi-scale feature encoding based on the dual attention mechanism includes:

[0013] The input low-light frame data and five-channel event voxels are mapped into a higher-dimensional embedding feature map through two-dimensional convolution;

[0014] The obtained embedded feature map is convolutionally downsampled to obtain the five-level feature map F of the low-light frame. i 0 Five-level feature map of event voxels Represents a series;

[0015] Layer normalization is performed on the input low-light frame level 5 feature map and event level 5 feature map. Channel-level local context aggregation is performed through convolution, and attention maps are calculated to output channel feature maps with enhanced channel information. and

[0016] Layer normalization is performed on the input low-light frame level 5 feature map and event level 5 feature map, cyclic shift operation is performed, and attention map is calculated to output a spatial feature map that deeply fuses global context information. and

[0017] A lightweight gating mechanism is used to dynamically generate channel feature map weights and spatial feature map weights. Adaptive fusion of channel and spatial features is achieved through weighted summation. Finally, the adaptively fused feature map F is... i 0 'and The input to the feedforward network yields the enhanced feature map F. i 1 and

[0018] Furthermore, the dynamically weighted feature fusion includes:

[0019] The obtained enhanced feature map is subjected to two residual feature enhancements to obtain the feature map F after preliminary enhancement. i 2 and

[0020] The enhanced feature map is then concatenated by channel stitching to obtain EF. i 0 EF is obtained by processing convolutional layers with different kernel sizes. i 1 The input is then fed into two parallel modules composed of convolutional layers to generate fusion weights ω for the channel-space bimodal features. i and EF i 1 Divide into frame features F along the channel dimension i 3 and event characteristics By weighted summation and inputting the result into the multi-scale feature encoding algorithm HAM based on the dual attention mechanism, a fused feature map EF is finally obtained, which integrates event edge information and frame color and texture information. i 3 .

[0021] Furthermore, image reconstruction and enhancement are performed by upsampling the fused feature maps to obtain the final clear image, including:

[0022] S11: Use the i-th level fusion feature map as the current level fusion feature map, i = 5;

[0023] S12: Upsample the current level fusion feature map to obtain an upsampled feature map. Then, concatenate the upsampled feature map with the (i-1)th level fusion feature map along the channel and pass it through the multi-scale feature coding algorithm HAM based on the dual attention mechanism to obtain the restored image.

[0024] S13: Use the restored image as the current level fusion feature map, and let i = i-1, repeat step S12 until i = 1; use the restored image obtained in the last iteration as the final clear image EF1. 4 .

[0025] Furthermore, the channel feature map weights are calculated using the following formula:

[0026]

[0027] Where Concat represents channel concatenation, ConLu represents 1x1 convolution with ReLU activation function, and Conv... 3×3This represents a 3x3 depth separable convolution, where σ represents the Sigmoid function. The channel feature map weights represent the values ​​of the low-light frames. The channel feature map weights represent the events.

[0028] Furthermore, the adaptive fusion feature map F i 0 'and Calculated using the following formula:

[0029]

[0030] Where ⊙ represents element-wise multiplication. The spatial feature map weights of low-light frames are represented. The weights of the spatial feature map representing the event.

[0031] Furthermore, the preliminarily enhanced feature map F i 2 and Calculated using the following formula:

[0032]

[0033] Here, Res represents the enhancement of residual features.

[0034] Furthermore, the EF is obtained through multi-kernel-size convolutional layers. i 1 Calculated using the following formula:

[0035]

[0036] Where GAP represents average pooling, Conv 1×1 2 This represents two sets of convolutional layer networks consisting of 1x1 convolutions, layer normalization, and ReLU activation functions. 3×3 2 This represents two sets of convolutional layer networks consisting of 3x3 convolutions, layer normalization, and ReLU activation functions. 5×5 2 This represents two sets of convolutional layer networks consisting of 5x5 convolutions, layer normalization, and ReLU activation functions. 1×1 σ represents a 1x1 convolution, and σ represents the Sigmoid function.

[0037] Furthermore, the fusion weight ω i and the fused feature map EF i 3 Calculated using the following formula:

[0038]

[0039] Among them, Conv 1×1 1 This represents a set of convolutional layer networks consisting of 1x1 convolutions, layer normalization, and ReLU activation functions. Indicates channel weight, Represents spatial weight, ω i Representing frame features F i 3 The weight.

[0040] Furthermore, the final clear image EF1 is obtained. 4 Calculated using the following formula:

[0041]

[0042] Among them, Conv 3×3 Represents a 3x3 convolution, Conv 1×1 This represents a 1x1 convolution.

[0043] The beneficial effects of this technical solution are:

[0044] This invention improves the fidelity of enhanced images: the event stream provides more accurate brightness change information, guiding the neural network to generate images that are closer to reality and have less noise. By fusing event stream data with image frame information, it effectively distinguishes between real image signals and noise, ensuring that the reconstructed image maintains a high degree of consistency with the real high-definition reference image at the pixel level.

[0045] It can improve the structural similarity of the enhanced image: the multi-scale feature encoding based on the dual attention mechanism and the dynamic weighted feature fusion mechanism enhance the edge sharpness and texture details. While improving pixel accuracy, it significantly enhances the overall structure, contour and texture preservation of the image, avoids excessive smoothing and structural distortion, and outputs a more natural and clear image.

[0046] It can optimize the perceived quality of enhanced images: event data supplements the information missing in traditional image frames in low-light scenes, suppresses artifacts and halos, and reconstructs details that are more consistent with advanced semantic perception. It achieves a high degree of consistency between enhanced images and real images at the semantic and perceptual levels, making them indistinguishable to the human eye and significantly improving the naturalness of visual perception.

[0047] The technical solution of this invention can adaptively weight and fuse multi-scale events and image features, effectively suppressing noise while restoring high-quality images with clear details and naturalness, significantly improving the perception ability of the visual system in low-light environments, and effectively solving the problems faced by existing solutions.

[0048] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description

[0049] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.

[0050] Figure 1 This is a flowchart illustrating an embodiment of the present invention;

[0051] Figure 2 This is an overall framework diagram of an embodiment of the present invention;

[0052] Figure 3 This is a schematic diagram of a multi-scale feature encoding algorithm based on a dual attention mechanism according to an embodiment of the present invention;

[0053] Figure 4 This is a schematic diagram of the attention-weighted fusion stage in an embodiment of the present invention. Detailed Implementation

[0054] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0055] A specific embodiment of the present invention discloses an event- and frame-based low-light image enhancement method, the flowchart of which is shown below. Figure 1 As shown, the method includes the following steps:

[0056] Simultaneously acquire event data generated by the event sensor in the event camera and low-light frame data generated by the APS frame sensor, and perform event stream preprocessing on the event data to convert the event data into five-channel event voxels;

[0057] Multi-scale feature encoding based on a dual attention mechanism is performed on the low-light frame data and the five-channel event voxels to enhance image features from both the channel dimension and the spatial dimension, thereby obtaining enhanced feature maps corresponding to the low-light frame data and event voxels.

[0058] Dynamically weighted feature fusion is performed on the enhanced feature maps corresponding to low-light frame data and event voxels. A convolution-based spatial-channel attention algorithm is used to adjust the weights of different modalities to obtain a fused feature map of the edge information of event voxels and the color and texture information of low-light frame data.

[0059] Image reconstruction and enhancement are performed by upsampling the fused feature map to obtain the final clear image.

[0060] In principle, an event camera is a novel image sensor based on biomimetic vision principles. Unlike traditional frame cameras that use a fixed frame rate for global sampling, its core functionality involves asynchronous data output through individual pixels responding independently to changes in light intensity. When a pixel receives a change in light intensity exceeding a preset threshold, it immediately triggers an event, outputting a data tuple containing pixel coordinates, a timestamp, and polarity (+1 for increased light intensity, -1 for decreased intensity). The core hardware advantage of the event camera lies in its integration of an APS frame sensor and an event sensor. These two sensors operate in parallel in low-light scenes, synchronously capturing different modalities of data from the same scene: the APS frame sensor generates low-light frame data, and the event sensor generates event data. Both sensors share the same optical system, and the event timestamp is precisely aligned with the exposure time window of the low-light frame, ensuring that the generated low-light frame data and event data correspond to "the same time period within the same low-light scene."

[0061] Specifically, the acquired data tuples cannot be directly input into the deep learning network; the event data needs to be converted into five-channel event voxels for subsequent processing by the deep learning network. For events e captured by the event camera... i =(x i ,y i ,t i ,p i ), (x i ,y i ) represents the pixel coordinates of the event, t i It is the timestamp of the last event triggered at that pixel, p i =±1 indicates the polarity of the event. To encode event data as a five-channel event voxel E, firstly, for a given group k of events... (This group contains N input events e) i ), calculate the time range ΔT = t for this group of events. N-1 -t0. Then, all event timestamps. Mapped to a time range [0,6] of length B=7:

[0062]

[0063] Finally, through linear interpolation, the polarity of events over the time range [0,6] is assigned to discrete integer timestamps t. n Construct an event voxel E on {0,1,…,6}:

[0064]

[0065] Among them, (xl ,y m ) represents the event coordinates of event voxel E. To suppress boundary noise, only t is retained. n Given ∈{1,2,…,5}, we obtain the final five-channel event voxel, represented as H represents the vertical coordinate range of the event voxel, and W represents the horizontal coordinate range of the event voxel.

[0066] When implementing, such as Figure 2 As shown in the feature encoding module of the overall framework diagram, the multi-scale feature encoding based on the dual attention mechanism includes:

[0067] The input low-light frame data and five-channel event voxels are mapped into a higher-dimensional embedding feature map through two-dimensional convolution;

[0068] The obtained embedded feature map is convolutionally downsampled to obtain the five-level feature map F of the low-light frame. i 0 Five-level feature map of event voxels Represents a series;

[0069] Layer normalization is performed on the input low-light frame level 5 feature map and event level 5 feature map. Channel-level local context aggregation is performed through convolution, and attention maps are calculated to output channel feature maps with enhanced channel information. and

[0070] Layer normalization is performed on the input low-light frame level 5 feature map and event level 5 feature map, cyclic shift operation is performed, and attention map is calculated to output a spatial feature map that deeply fuses global context information. and

[0071] A lightweight gating mechanism is used to dynamically generate channel feature map weights and spatial feature map weights. Adaptive fusion of channel and spatial features is achieved through weighted summation. Finally, the adaptively fused feature map F is... i 0 'and The input to the feedforward network yields the enhanced feature map F. i 1 and

[0072] Specifically, the schematic diagram of the multi-scale feature encoding algorithm based on the dual attention mechanism is as follows: Figure 3As shown, to capture richer and more complex semantic features, the data embedding algorithm maps the input low-light frame F and the five-channel event voxel E into higher-dimensional embedding feature maps F' and E' through two-dimensional convolution. Then, to obtain multi-scale features at different resolutions, convolutional downsampling is performed on the embedding feature maps F' and E' to obtain five-level feature maps F' and E' of the low-light frame and event voxel. i 0 , i∈{1,2,…,5} represents the level, and the superscript 0 indicates the stage number. The size of the five-level feature map for low-light frames and event voxels is...

[0073] Furthermore, for the i-th layer feature pyramid, the low-light frame feature map F i 0 and event feature map Input two parallel channel attention algorithms and a spatial attention algorithm respectively.

[0074] The channel attention algorithm includes: first, processing the input low-light frame feature map F... i 0 and event feature map Layer normalization is performed; subsequently, channel-level local context aggregation is performed using a 1x1 convolution. Next, a 3x3 convolution is used to generate the query vector Q, key vector K, and value vector V. Then, a self-attention mechanism is used to calculate the cross-covariance between feature channel dimensions, and the attention map is calculated according to the following formula:

[0075]

[0076] Here, d represents the dimensions of the key vector K and the query vector Q, and Softmax represents the soft maximum normalization function. This attention map effectively captures the channel dependencies of global features. Finally, the attention map is transformed and fused using a 1x1 convolution, ultimately outputting a channel feature map with enhanced channel information. and

[0077] Spatial attention algorithms include: input low-light frame feature map F i 0 and event feature map First, the feature map undergoes layer normalization, followed by a cyclic shift operation, moving it 4 pixels to the right and 4 pixels down. After deformation, the shifted feature map is processed through a linear layer to generate a query vector Q, a key vector K, and a value vector V. An attention map is then calculated. Finally, the map is processed through deformation and a linear layer to output a spatial feature map that deeply integrates global contextual information while preserving local details. and

[0078] Next, an adaptive gating network is used to dynamically generate channel feature map weights ω using a lightweight gating mechanism. i Spatial feature map weights 1-ω i This enables adaptive extraction and fusion of channel features and spatial features.

[0079] The network first uses the low-light frame feature map F i 0 and event feature map Enhanced channel feature map of channel information and Spatial feature maps that incorporate global context information and The features are concatenated along the channel dimension to form rich contextual information. Then, a 1x1 convolution and ReLU activation function are used to compress the number of channels. Next, a 3x3 depthwise separable convolution is used to capture local spatial contextual relationships, significantly reducing the number of parameters while maintaining a sufficient receptive field. Finally, a 1x1 convolution is used to compress the feature channels to a single channel, and the channel feature map weights are generated using the Sigmoid function.

[0080] Specifically, the channel feature map weights are calculated using the following formula:

[0081]

[0082] Where Concat represents channel concatenation, ConLu represents 1x1 convolution with ReLU activation function, and Conv... 3×3 This represents a 3x3 depth separable convolution, where σ represents the Sigmoid function. The channel feature map weights represent the values ​​of the low-light frames. The channel feature map weights represent the events.

[0083] Obtain the channel feature map weights ω i Then, the spatial feature map weights can be expressed as 1-ω i Based on the channel feature map weights and spatial feature map weights, adaptive extraction is achieved through weighted summation, and F is calculated. i 0 'and

[0084] Specifically, the adaptive fusion feature map F i 0 'and Calculated using the following formula:

[0085]

[0086] Where ⊙ represents element-wise multiplication. The spatial feature map weights of low-light frames are represented. The weights of the spatial feature map representing the event.

[0087] Finally, F i 0 'and The input to the feedforward network further enhances the representation of features, resulting in feature maps F that encode channel and spatial information. i 1 and F i 0 'and First, the features undergo layer normalization, followed by a 1x1 convolution to expand the channel dimension. Then, a 3x3 convolution encodes spatial local context information. Afterward, the features from the 3x3 convolution are split in two along the channel dimension, with one part introduced into a non-linearity using a Gaussian Error Linear Unit (GELU) activation function. Finally, the two parts are multiplied element-wise to enhance the features. Finally, a 1x1 convolution compresses the number of channels back to the original dimension and is then compared with the original input F of the module. i 0 'and Perform residual connections to output an enhanced feature F that contains both channel and spatial information. i 1 and

[0088] When implementing, such as Figure 2 As shown in the feature fusion module of the overall framework diagram, the dynamically weighted feature fusion includes:

[0089] The obtained enhanced feature map is subjected to two residual feature enhancements to obtain the feature map F after preliminary enhancement. i 2 and

[0090] The enhanced feature map is then concatenated by channel stitching to obtain EF. i 0 EF is obtained by processing convolutional layers with different kernel sizes. i 1 The input is then fed into two parallel modules composed of convolutional layers to generate fusion weights ω for the channel-space bimodal features. i and EF i 1 Divide into frame features F along the channel dimension i 3 and event characteristics By weighted summation and inputting the result into the multi-scale feature encoding algorithm HAM based on the dual attention mechanism, a fused feature map EF is finally obtained, which integrates event edge information and frame color and texture information.i 3 .

[0091] Specifically, the dynamically weighted feature fusion includes a residual fusion stage and an attention-weighted fusion stage;

[0092] The residual fusion stage includes: processing feature maps F i 1 and Perform two residual feature enhancements to obtain F. i 2 ,

[0093]

[0094] Furthermore, the preliminarily enhanced feature map F i 2 and Calculated using the following formula:

[0095]

[0096] Here, Res represents the enhancement of residual features.

[0097] The attention-weighted fusion stage includes: (See diagram below) Figure 4 As shown, the frame features F obtained in the residual fusion stage are first... i 2 and event characteristics EF is obtained by splicing along the channel dimension i 0 This preserves information from both modalities. Then, the concatenated feature EF... i 0 The results were obtained by passing convolutional layers (composed of pooling, convolution, layer normalization, and ReLU activation function) with kernel sizes of 1, 3, and 5 respectively. The three components are concatenated along the channel, then subjected to a 1x1 convolution and a Sigmoid function, and finally connected to EF. i 0 Multiply the elements together, then multiply by EF. i 0 Residual join, to obtain EF i 1 .

[0098] Specifically, the EF is obtained through multi-kernel-size convolutional layers. i 1 Calculated using the following formula:

[0099]

[0100]

[0101] Where GAP represents average pooling, Conv 1×1 2 This represents two sets of convolutional layer networks consisting of 1x1 convolutions, layer normalization, and ReLU activation functions. 3×3 2 This represents two sets of convolutional layer networks consisting of 3x3 convolutions, layer normalization, and ReLU activation functions. 5×5 2 This represents two sets of convolutional layer networks consisting of 5x5 convolutions, layer normalization, and ReLU activation functions. 1×1 σ represents a 1x1 convolution, and σ represents the Sigmoid function.

[0102] Next, EF i 1 The input is then fed into two parallel modules composed of convolutional layers to generate fusion weights for frame features, and the EF is then... i 1 Divide into frame features F along the channel dimension i 3 and event characteristics

[0103]

[0104] In this context, split(·) represents dividing along the channel dimension.

[0105] Since event features and frame features are complementary, the generated weights ω i Used for frame feature F i 3 and event characteristics The fusion weights can be expressed as 1-ω i Weighted calculation of the fused feature map EF i 2 Finally, the fused feature map EF i 2 The high-level semantic fusion feature EF is further extracted using the HAM multi-scale feature encoding algorithm based on the dual attention mechanism. i 3 .

[0106] Specifically, the fusion weight ω i and the fused feature map EF i 3 Calculated using the following formula:

[0107]

[0108] Among them, Conv 1×1 1This represents a set of convolutional layer networks consisting of 1x1 convolutions, layer normalization, and ReLU activation functions. Indicates channel weight, Represents spatial weight, ω i Representing frame features F i 3 The weight.

[0109] When implementing, such as Figure 2 As shown in the overall framework diagram, the image enhancement module performs image reconstruction and enhancement by upsampling the fused feature map to obtain the final clear image, including:

[0110] S11: Use the i-th level fusion feature map as the current level fusion feature map, i = 5;

[0111] S12: Upsample the current level fusion feature map to obtain an upsampled feature map. Then, concatenate the upsampled feature map with the (i-1)th level fusion feature map along the channel and pass it through the multi-scale feature coding algorithm HAM based on the dual attention mechanism to obtain the restored image.

[0112] S13: Use the restored image as the current level fusion feature map, and let i = i-1, repeat step S12 until i = 1; use the restored image obtained in the last iteration as the final clear image EF1. 4 .

[0113] Specifically, EF is decoded step by step during the image reconstruction stage. i 3 Features are extracted and gradually fused with the previous layer to restore the spatial resolution of the image. By fusing coarse and fine features, it is ensured that global contextual information is preserved while local details are meticulously restored during image reconstruction.

[0114] Furthermore, the final clear image EF1 is obtained. 4 Calculated using the following formula:

[0115]

[0116] Among them, Conv 3×3 Represents a 3x3 convolution, Conv 1×1 This represents a 1x1 convolution.

[0117] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.

[0118] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A low-light image enhancement method based on events and frames, characterized in that, The method includes the following steps: Simultaneously acquire event data generated by the event sensor in the event camera and low-light frame data generated by the APS frame sensor, and perform event stream preprocessing on the event data to convert the event data into five-channel event voxels; Multi-scale feature encoding based on a dual attention mechanism is performed on the low-light frame data and the five-channel event voxels to enhance image features from both the channel dimension and the spatial dimension, thereby obtaining enhanced feature maps corresponding to the low-light frame data and event voxels. Dynamically weighted feature fusion is performed on the enhanced feature maps corresponding to low-light frame data and event voxels. A convolution-based spatial-channel attention algorithm is used to adjust the weights of different modalities to obtain a fused feature map of the edge information of event voxels and the color and texture information of low-light frame data. Image reconstruction and enhancement are performed by upsampling the fused feature map to obtain the final clear image.

2. The low-light image enhancement method based on events and frames according to claim 1, characterized in that, The multi-scale feature encoding based on the dual attention mechanism includes: The input low-light frame data and five-channel event voxels are mapped into a higher-dimensional embedding feature map through two-dimensional convolution; The obtained embedded feature map is convolutionally downsampled to obtain the five-level feature map F of the low-light frame. i 0 Five-level feature map of event voxels Represents a series; Layer normalization is performed on the input low-light frame level 5 feature map and event level 5 feature map. Channel-level local context aggregation is performed through convolution, and attention maps are calculated to output channel feature maps with enhanced channel information. and Layer normalization is performed on the input low-light frame level 5 feature map and event level 5 feature map, cyclic shift operation is performed, and attention map is calculated to output a spatial feature map that deeply fuses global context information. and A lightweight gating mechanism is used to dynamically generate channel feature map weights and spatial feature map weights. Adaptive fusion of channel and spatial features is achieved through weighted summation. Finally, the adaptively fused feature map F is... i 0 'and The input to the feedforward network yields the enhanced feature map F. i 1 and 3. The low-light image enhancement method based on events and frames according to claim 1, characterized in that, The dynamically weighted feature fusion includes: The obtained enhanced feature map is subjected to two residual feature enhancements to obtain the feature map F after preliminary enhancement. i 2 and The enhanced feature map is then concatenated by channel stitching to obtain EF. i 0 EF is obtained by processing convolutional layers with different kernel sizes. i 1 The input is then fed into two parallel modules composed of convolutional layers to generate fusion weights ω for the channel-space bimodal features. i and EF i 1 Divide into frame features F along the channel dimension i 3 and event characteristics By weighted summation and inputting the result into the multi-scale feature encoding algorithm HAM based on the dual attention mechanism, a fused feature map EF is finally obtained, which integrates event edge information and frame color and texture information. i 3 .

4. The low-light image enhancement method based on events and frames according to claim 1, characterized in that, Image reconstruction and enhancement by upsampling the fused feature map yields a final, clear image, including: S11: Use the i-th level fusion feature map as the current level fusion feature map, i = 5; S12: Upsample the current level fusion feature map to obtain an upsampled feature map. Then, concatenate the upsampled feature map with the (i-1)th level fusion feature map along the channel and pass it through the multi-scale feature coding algorithm HAM based on the dual attention mechanism to obtain the restored image. S13: Use the restored image as the current level fusion feature map, and let i = i-1, repeat step S12 until i = 1; use the restored image obtained in the last iteration as the final clear image EF1. 4 .

5. The low-light image enhancement method based on events and frames according to claim 2, characterized in that, The channel feature map weights are calculated using the following formula: Where Concat represents channel concatenation, ConLu represents 1x1 convolution with ReLU activation function, and Conv... 3×3 This represents a 3x3 depth separable convolution, where σ represents the Sigmoid function. The channel feature map weights represent the values ​​of the low-light frames. The channel feature map weights represent the events.

6. The low-light image enhancement method based on events and frames according to claim 2, characterized in that, The adaptive fusion feature map F i 0 'and Calculated using the following formula: Where ⊙ represents element-wise multiplication. The spatial feature map weights of low-light frames are represented. The weights of the spatial feature map representing the event.

7. The low-light image enhancement method based on events and frames as described in claim 3, characterized in that, The pre-enhanced feature map F i 2 and Calculated using the following formula: Here, Res represents the enhancement of residual features.

8. The low-light image enhancement method based on events and frames according to claim 3, characterized in that, The EF is obtained by processing convolutional layers with multiple kernel sizes. i 1 Calculated using the following formula: Where GAP represents average pooling, Conv 1×1 2 This represents two sets of convolutional layer networks consisting of 1x1 convolutions, layer normalization, and ReLU activation functions. 3×3 2 This represents two sets of convolutional layer networks consisting of 3x3 convolutions, layer normalization, and ReLU activation functions. 5×5 2 This represents two sets of convolutional layer networks consisting of 5x5 convolutions, layer normalization, and ReLU activation functions. 1×1 σ represents a 1x1 convolution, and σ represents the Sigmoid function.

9. The low-light image enhancement method based on events and frames according to claim 3, characterized in that, The fusion weight ω i and the fused feature map EF i 3 Calculated using the following formula: Among them, Conv 1×1 1 This represents a set of convolutional layer networks consisting of 1x1 convolutions, layer normalization, and ReLU activation functions. Indicates channel weight, Represents spatial weight, ω i Representing frame features F i 3 The weight.

10. The low-light image enhancement method based on events and frames according to claim 4, characterized in that, The final clear image EF1 is obtained. 4 Calculated using the following formula: Among them, Conv 3×3 Represents a 3x3 convolution, Conv 1×1 This represents a 1x1 convolution.