Event-based low-light image enhancement method and electronic device

By acquiring event information and utilizing a pre-trained low-light image enhancement model, enhancement processing is performed on both the image and event feature maps, solving the problem of detail loss in extremely dark areas and improving the video enhancement effect.

CN118096626BActive Publication Date: 2026-08-25BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410165178.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-05
Publication Date
2026-08-25
Estimated Expiration
2044-02-05

AI Technical Summary

Technical Problem

Existing technologies for video enhancement in extremely dark areas result in the loss of detail and a decline in visual quality, making effective recovery difficult.

Method used

By acquiring event information from the initial image data, a pre-trained low-light image enhancement model, including a reflection component enhancement module, an illumination component enhancement module, and a synthesis module, is used to map the image feature map and event feature map respectively, enhance the illumination component and reflection component, and finally synthesize the enhanced image data.

Benefits of technology

Restoring detail in extremely dark areas improves video enhancement and enhances visual quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118096626B_ABST
    Figure CN118096626B_ABST
Patent Text Reader

Abstract

The present disclosure provides an event-based low-light image enhancement method and an electronic device. The initial image data is obtained, and the event information corresponding to the initial image data is determined. The initial image data and the event information are input into a pre-trained low-light image enhancement model. The initial image data and the event information are respectively mapped to obtain an image feature map and an event feature map. The image feature map is input into an illumination component enhancement module, and the image feature map is enhanced by the illumination component enhancement module to obtain an illumination component. The image feature map and the event feature map are input into a reflection component enhancement module, and the image feature map is enhanced based on the event feature map by the reflection component enhancement module to obtain a reflection component. The illumination component and the reflection component are input into a synthesis module, and the synthesis module is used for synthesis processing to obtain enhanced image data. The details of the extremely dark area of the image are enhanced, and the enhancement effect of the low-light image enhancement is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing, and more particularly to an event-based method and electronic device for enhancing low-light images. Background Technology

[0002] Limited by low-light environments and camera sensor limitations, videos often suffer from significant noise and insufficient information, leading to loss of detail and a decline in overall visual quality. Enhancement algorithms can significantly improve the details of low-light videos, enhancing their visual quality and applications in downstream tasks such as autonomous driving, object detection, and semantic segmentation.

[0003] Current video enhancement methods typically extract and enhance information from the video itself. However, due to the limitations of camera sensors, extremely dark areas of the video inevitably lose detail, which is still lost during enhancement, resulting in a decrease in quality.

[0004] Therefore, how to enhance the details in extremely dark areas of a video has become an important technical problem. Summary of the Invention

[0005] In view of this, the purpose of this disclosure is to propose an event-based low-light image enhancement method and electronic device to solve or partially solve the above problems.

[0006] To achieve the above objectives, a first aspect of this disclosure provides an event-based low-light image enhancement method, the method comprising:

[0007] Acquire initial image data and determine the event information corresponding to the initial image data;

[0008] The initial image data and the event information are input into a pre-trained low-light image enhancement model, wherein the low-light image enhancement model includes a reflection component enhancement module, an illumination component enhancement module, and a synthesis module;

[0009] The initial image data and the event information are respectively mapped to obtain image feature maps and event feature maps;

[0010] The image feature map is input to the illumination component enhancement module, and the illumination component enhancement module is used to enhance the image feature map to obtain the illumination component.

[0011] The image feature map and the event feature map are input into the reflection component enhancement module. The reflection component enhancement module enhances the image feature map based on the event feature map to obtain the reflection component.

[0012] The illumination component and the reflection component are input into the synthesis module, and the synthesis module is used to perform synthesis processing to obtain enhanced image data.

[0013] Based on the same inventive concept, a second aspect of this disclosure proposes an event-based low-light image enhancement device, comprising:

[0014] The data acquisition module is configured to acquire initial image data and determine the event information corresponding to the initial image data;

[0015] The data input module is configured to input the initial image data and the event information into a pre-trained low-light image enhancement model, wherein the low-light image enhancement model includes a reflection component enhancement module, an illumination component enhancement module, and a synthesis module;

[0016] The mapping processing module is configured to perform mapping processing on the initial image data and the event information respectively to obtain image feature maps and event feature maps;

[0017] The illumination component enhancement module is configured to input the image feature map into the illumination component enhancement module, and use the illumination component enhancement module to enhance the image feature map to obtain the illumination component.

[0018] The reflection component enhancement module is configured to input the image feature map and the event feature map into the reflection component enhancement module, and use the reflection component enhancement module to enhance the image feature map based on the event feature map to obtain the reflection component;

[0019] The compositing module is configured to input the illumination component and the reflection component into the compositing module, and use the compositing module to perform compositing processing to obtain enhanced image data.

[0020] Based on the same inventive concept, a third aspect of this disclosure proposes an electronic device including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor, when executing the computer program, implements the event-based low-light image enhancement method as described above.

[0021] Based on the same inventive concept, a fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the event-based low-light image enhancement method as described above.

[0022] As can be seen from the above, this disclosure proposes an event-based low-light image enhancement method and electronic device. It acquires initial image data and determines the event information corresponding to the initial image data. The event information contains complete motion information corresponding to the initial image data. The initial image information and event information are input into a pre-trained low-light image enhancement model for subsequent output of enhanced image data. The initial image data and the event information are mapped to obtain image feature maps and event feature maps, which are then input into an illumination component enhancement module and a reflection component enhancement module for enhancement processing. The image feature maps are input into the illumination component enhancement module, which enhances the image feature maps to obtain the illumination component. The image feature maps and event feature maps are input into the reflection component enhancement module, which enhances the image feature maps based on the event feature maps to obtain the reflection component. The illumination component and reflection component are input into a synthesis module, which performs synthesis processing to obtain enhanced image data. By processing the initial image data and temporal information using the aforementioned low-light image enhancement model, the resulting enhanced image data is more accurate. Furthermore, the use of event information during processing allows for the recovery of details in extremely dark areas, thus improving the enhancement effect. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart of an event-based low-light image enhancement method according to an embodiment of this disclosure;

[0025] Figure 2 This is a structural block diagram of an event-based low-light image enhancement device according to an embodiment of the present disclosure;

[0026] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0028] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0029] The following are definitions of terms used in this disclosure:

[0030] Retinex: Retinex is a commonly used image enhancement method based on scientific experiments and analysis. Retinex is based on the theory of the retina and cerebral cortex.

[0031] U-Net Network: U-Net is one of the earliest algorithms to use fully convolutional networks for semantic segmentation. The paper used a symmetrical U-shaped structure containing compression and expansion paths, which was very innovative at the time and influenced the design of several subsequent segmentation networks to some extent. The name of the network is also derived from its U-shaped shape.

[0032] The SDSD dataset (Seeing Dynamic Scenes in the Dark) was proposed in the paper Seeing Dynamic Scene in the Dark: A High-Quality Video Dataset with Mechatronic Alignment. The SDSD dataset contains low-light videos and their paired normal-light videos. The videos were captured by a Canon EOS 6D Mark II camera and contain 70 indoor video pairs and 80 outdoor video pairs.

[0033] V2E: Event Camera Emulator.

[0034] DSEC Dataset: DSEC is a stereo camera dataset for driving scenarios, containing data from two monochrome event cameras and two global shutter color cameras.

[0035] AdamW Optimizer: The Adam optimizer combines the advantages of both AdaGrad and RMSProp optimization algorithms. It comprehensively considers the first moment estimation (mean of the gradient) and the second moment estimation (uncentered variance of the gradient) to calculate the update step size.

[0036] Based on the above description, this embodiment proposes an event-based low-light image enhancement method, such as... Figure 1 As shown, the method includes:

[0037] Step 101: Obtain initial image data and determine the event information corresponding to the initial image data.

[0038] In specific implementation, initial image data is acquired, wherein the initial image data is low-light image data that needs to be enhanced, and the image data includes at least one of the following forms: video or image.

[0039] The event information corresponding to the initial image data is determined. This event information is the content captured by the event camera, and the captured content is identical to the content corresponding to the initial image data. The event camera is a novel sensor characterized by its large dynamic range and fast response. Events captured by the event camera can completely record motion information even in extremely dark areas, making it ideal for low-light video enhancement.

[0040] In some embodiments, the event information may be information that has been captured in advance and stored in a database. When storing, the event information is associated with the corresponding image data and stored together. Subsequently, the event information corresponding to the initial image information can be retrieved from the database based on the initial image information.

[0041] In some embodiments, after acquiring initial image data, the initial image data is sent to the event camera, which then obtains event information based on the initial image data. The event information sent by the event camera is received to obtain the event information corresponding to the initial image data.

[0042] Step 102: Input the initial image data and the event information into a pre-trained low-light image enhancement model, wherein the low-light image enhancement model includes a reflection component enhancement module, an illumination component enhancement module, and a synthesis module.

[0043] In practice, the initial low-light image enhancement model is trained in advance to obtain the low-light image enhancement model. The input of the low-light image enhancement model is the initial image data and event information, and the output is the enhanced image data.

[0044] The initial image data and event information are input into the low-light image enhancement model, which includes a reflection component enhancement module, an illumination component enhancement module, and a synthesis module, so that the low-light image enhancement model can process the initial image data and event information to obtain enhanced image data.

[0045] Step 103: Map the initial image data and the event information respectively to obtain image feature map and event feature map.

[0046] Step 104: Input the image feature map into the illumination component enhancement module, and use the illumination component enhancement module to enhance the image feature map to obtain the illumination component.

[0047] Step 105: Input the image feature map and the event feature map into the reflection component enhancement module, and use the reflection component enhancement module to enhance the image feature map based on the event feature map to obtain the reflection component.

[0048] In specific implementation, the initial image data is mapped to the coding space to obtain an image feature map, and the event information is mapped to the coding space to obtain an event feature map. The obtained image feature maps are input to the illumination component enhancement module, which performs enhancement processing to obtain the illumination component. The obtained image feature map and event feature map are input to the reflection component enhancement module, which enhances the image feature map based on the event feature map to obtain the reflection component.

[0049] Step 106: Input the illumination component and the reflection component into the synthesis module, and use the synthesis module to perform synthesis processing to obtain enhanced image data.

[0050] In specific implementation, the enhanced illumination component and reflection component are input into the synthesis module, and the illumination component and reflection component are synthesized to obtain enhanced image data, which is normal light image data.

[0051] The above scheme involves acquiring initial image data and determining the corresponding event information, where the event information contains complete motion information. The initial image information and event information are input into a pre-trained low-light image enhancement model for subsequent output of enhanced image data. The initial image data and event information are mapped to obtain image feature maps and event feature maps, which are then input into the illumination component enhancement module and reflection component enhancement module for enhancement. The image feature maps are input into the illumination component enhancement module, which enhances them to obtain the illumination component. The image feature maps and event feature maps are input into the reflection component enhancement module, which enhances them based on the event feature maps to obtain the reflection component. Finally, the illumination component and reflection component are input into a synthesis module for synthesis to obtain the enhanced image data. By processing the initial image data and temporal information using the aforementioned low-light image enhancement model, the resulting enhanced image data is more accurate. Furthermore, the use of event information during processing allows for the recovery of details in extremely dark areas, thus improving the enhancement effect.

[0052] In some embodiments, step 103 specifically includes:

[0053] Step 1031: Obtain a preset residual block, input the initial image data and the event information into the residual block, and process the residual block to obtain the initial image feature map and the initial event feature map.

[0054] In practice, a preset residual block is obtained, and the initial image data and the event information are input into the residual block. The residual block is then mapped to the coding space to obtain the initial image feature map and the initial event feature map. Both the initial image feature map and the initial event feature map are feature maps of a fixed scale.

[0055] In some embodiments, the residual block consists of four convolutional blocks with 32, 16, 16, and 3 intermediate channels, respectively.

[0056] Step 1032: Overlay the initial image data and the initial image feature map to obtain the image feature map.

[0057] Step 1033: Overlay the event information and the initial event feature map to obtain the event feature map.

[0058] In specific implementation, the initial image data is superimposed with the mapped initial image feature map to obtain an image feature map. The event information is superimposed with the mapped initial event feature map to obtain an event feature map.

[0059] The above scheme, after obtaining the initial image feature map and the initial time feature map, overlays the initial image data and event information respectively, so that the initial image feature map contains the initial image features corresponding to the initial image data, and the initial event feature map contains the initial event features corresponding to the event information, so as to improve the enhancement effect when using the low-light image enhancement model for subsequent enhancement.

[0060] In some embodiments, the illumination component enhancement module includes an encoder, and step 104 specifically includes:

[0061] Step 1041: Use the encoder to perform dimensionality reduction processing on the image feature map to obtain the target image feature map.

[0062] Step 1042: Component extraction is performed on the feature map of the target image to obtain the initial illumination component.

[0063] Step 1043: Enhance the initial illumination component to obtain the illumination component.

[0064] In a specific implementation, the illumination component enhancement module includes an encoder. The image feature map is input into the encoder in the illumination component enhancement module, and the encoder is used to perform dimensionality reduction processing on the image feature map to obtain the target image feature map.

[0065] In some embodiments, the encoder is a pre-trained VQGAN encoder to transform high-dimensional features into a low-dimensional spatial representation.

[0066] The target image feature map is subjected to component extraction to obtain the initial illumination component. The initial illumination component is then enhanced to obtain the final illumination component. In this embodiment, U-Net is used for the extraction and enhancement of the initial illumination component. That is, U-Net is used to extract the initial illumination component from the target image feature map and enhance it to obtain the enhanced illumination component.

[0067] The above scheme reduces the complexity of the data and improves the computational speed of data processing by using an encoder for dimensionality reduction.

[0068] In some embodiments, the reflection component enhancement module includes an image feature processing module and an event feature processing module, and step 105 specifically includes:

[0069] Step 1051: Input the event feature map into the event feature processing module, and obtain the target event features through the event feature processing module.

[0070] Step 1052: Input the image feature map and the target event feature into the image feature processing module, and use the target event feature to enhance the image feature map to obtain the reflection component.

[0071] In specific implementation, the reflection component enhancement module includes an image feature processing module and an event feature processing module. The event feature map is input to the event feature processing module, which extracts features from the event feature map to obtain target event features. The image feature map and the extracted target event features are then input to the image feature processing module, which enhances the image feature map using the target event features to obtain the reflection component.

[0072] In some embodiments, the image feature processing module includes a feature extraction module, a downsampling layer, and an image event fusion module, and step 1052 specifically includes:

[0073] Step 10521: Input the image feature map into the feature extraction module, and use the feature extraction module to extract features from the image feature map to obtain initial image features.

[0074] Step 10522: Input the initial image features into the downsampling layer, and use the downsampling layer to amplify the number of channels of the initial image features to obtain image features.

[0075] Step 10523: Input the image features and the target event features into the image event fusion module, and use the target event features to enhance the image features to obtain the reflection component.

[0076] In specific implementation, the image feature processing module includes a feature extraction module, a downsampling layer, and an image event fusion module. The image feature map is input to the feature extraction module, which extracts features from the image feature map to obtain initial image features, which are then input to the downsampling layer.

[0077] The downsampling layer is a convolutional layer that amplifies the number of channels in the initial image features. After passing through the downsampling layer, the number of channels in the initial image features becomes twice that before being input to the downsampling layer.

[0078] The extracted image features and target event features are input into the image event fusion module, and the image features are enhanced using the target event features to obtain the reflection component.

[0079] In some embodiments, the feature extraction module includes four layers, each layer including a first normalization layer, a first convolutional layer, a channel attention layer, a second normalization layer, and a second convolutional layer. Step 10521 specifically includes:

[0080] The image feature map is input into the first layer, processed by the first layer to obtain the processing result, and the processing result is input into the next layer. This process is repeated until the image is input into the fourth layer, processed by the fourth layer, and the resulting processing result is used as the initial image feature.

[0081] The specific processing steps at each layer include:

[0082] The image feature map is input into the first normalization layer, processed by the first normalization layer to obtain the first image feature, and the first image feature is sent to the first convolutional layer.

[0083] The first image features are processed using the first convolutional layer to obtain the second image features;

[0084] The second image feature is input into the channel attention layer, and the third image feature is obtained through the channel attention layer.

[0085] The image feature map is superimposed with the third image feature to obtain the fourth image feature, and the fourth image feature is sent to the second normalization layer;

[0086] The fourth image feature is processed using the second normalization layer to obtain the fifth image feature;

[0087] The fifth image feature is input into the second convolutional layer, and the sixth image feature is output after processing by the second convolutional layer.

[0088] The sixth image feature and the fourth image feature are superimposed to obtain the processing result.

[0089] In specific implementation, each layer in the feature extraction module includes a first normalization layer, a first convolutional layer, a channel attention layer, a second normalization layer, and a second convolutional layer, and they are connected in the order of the first normalization layer, the first convolutional layer, the channel attention layer, the second normalization layer, and the second convolutional layer.

[0090] The specific process of inputting the image feature map into the first layer and iteratively inputting it into the fourth layer is as follows:

[0091] Image features are input into the first layer, processed by the first layer, and the processing result is input into the second layer for further processing. The processing result obtained from the second layer is input into the third layer for further processing, and the processing result obtained from the third layer is input into the fourth layer for further processing. The processing result obtained from the fourth layer is used as the initial image features.

[0092] The specific processing steps at each layer include:

[0093] The image feature map is input to a first normalization layer, processed by the first normalization layer to obtain a first image feature, and then sent to a first convolutional layer. The first convolutional layer processes the first image feature to obtain a second image feature. The second image feature is then input to a channel attention layer, processed by the channel attention layer to obtain a third image feature.

[0094] A skip connection is added between the first normalization layer and the channel attention layer, that is, the image feature map is superimposed with the third image feature to obtain the fourth image feature, and the fourth image feature is sent to the second normalization layer.

[0095] The fourth image feature is processed by the second normalization layer to obtain the fifth image feature. The fifth image feature is then input into the second convolutional layer, and the sixth image feature is output after processing by the second convolutional layer.

[0096] A skip connection is added between the second normalization layer and the convolutional layer, that is, the sixth image feature and the fourth image feature are superimposed to obtain the processing result.

[0097] In some embodiments, the image event fusion module includes a region selection module and an attention fusion module, and step 10523 specifically includes:

[0098] Step 105231: Input the image features into the region selection module to determine the image information of multiple image regions corresponding to the image features.

[0099] Step 105232: Select the image region corresponding to the image information that meets the first preset condition as the target image region, and send the target image region to the attention fusion module.

[0100] In practice, image features are input to the region selection module to obtain image information of multiple image regions corresponding to the image features. The image information includes brightness information and signal-to-noise ratio information. The image information of each image region is compared with a first preset condition, and the image region corresponding to the image information that meets the first preset condition is selected as the target image region. The target image region is then sent to the attention fusion module, and the target image region is the extremely dark region.

[0101] In some embodiments, the first preset condition is that the brightness is less than a preset brightness threshold and the signal-to-noise ratio is less than a preset signal-to-noise ratio threshold.

[0102] Step 105233: In the attention fusion module, the target image region is enhanced using target event features to obtain the initial reflection component.

[0103] Step 105234: The initial reflection component is superimposed with the initial image features to obtain the reflection component.

[0104] In practice, the attention fusion module utilizes a cross-modal attention mechanism to enhance the selected target image region using the target event features, thereby obtaining an initial reflection component. The attention fusion module employs skip connections before and after the initial reflection component, that is, it superimposes the initial reflection component with the initial image features to obtain the final reflection component.

[0105] The attention fusion module is expressed by the following formula:

[0106]

[0107] Where F is the reflection component, The target image region is obtained by passing the image features of extremely dark areas through a convolutional layer with a kernel of 1. id represents the image features, and K... e V is obtained by passing the target event features through a first convolutional layer with a kernel of 1. e The target event features are obtained by passing them through a second convolutional layer with a kernel of 1. The parameters of the first and second convolutional layers are the same, e is the target event feature, and c is the number of channels of the image features.

[0108] The above scheme compares the image information of an image region with a first preset condition, identifying the image region that meets the first preset condition as the target image region, i.e., the extremely dark region. Subsequent enhancement only needs to be applied to the extremely dark region based on the characteristics of the target event, avoiding unnecessary enhancement operations on regions that do not require enhancement.

[0109] In some embodiments, the low-light image enhancement model further includes an alignment module, and before step 106, it further includes:

[0110] Step 10A: Input the illumination component and the reflection component into the alignment module to determine the first channel dimension corresponding to the illumination component and the second channel dimension corresponding to the reflection component.

[0111] Step 10B: Based on the first channel dimension and the second channel dimension, the illumination component and the reflection component are spliced ​​together to obtain the target feature vector.

[0112] In specific implementation, the low-light image enhancement model further includes an alignment module, which inputs the illumination component and the reflection component into the alignment module for alignment. The alignment module is a convolutional layer, meaning that the illumination component and the reflection component are input into the convolutional layer for alignment.

[0113] The alignment process is as follows:

[0114] Obtain the first channel dimension corresponding to the illumination component and the second channel dimension corresponding to the reflection component. Based on the first channel dimension and the second channel dimension, perform a splicing process on the illumination component and the reflection component to obtain a target feature vector. The target feature vector is the feature vector obtained after alignment.

[0115] In some embodiments, step 106 specifically includes:

[0116] Step 1061: Input the target feature vector into the synthesis module, wherein the synthesis module comprises four layers.

[0117] Step 1062: Using the target feature vector as candidate processing data, and taking the first layer of the synthesis module as the target layer number, iteratively process the candidate processing data. The specific process of each iteration is as follows:

[0118] Obtain the processing result of the feature extraction module corresponding to the target layer, superimpose the target feature vector with the processing result, and use it as updated candidate processing data. Then, take the next layer of the target layer as the target layer of the next iteration according to the layer of the synthesis module.

[0119] The iteration operation continues until the target layer does not have a next layer according to the layer number of the synthesis module;

[0120] The candidate processed data after iterative operations are used as the enhanced image data.

[0121] In specific implementation, the synthesis module includes four layers, each consisting of a feature extraction unit and an upsampling layer. The feature extraction unit is the same as the feature extraction module in the reflection enhancement processing module, and the upsampling layer consists of a convolutional layer and a pixelShuffle layer.

[0122] Using the target feature vector as candidate processing data, and taking the first layer of the synthesis module as the target layer number, the candidate processing data is iteratively processed. The specific process of each iteration is as follows:

[0123] A skip connection is added between the image feature processing module of the synthesis module and the reflection enhancement processing module to obtain the processing result of the feature extraction module corresponding to the target layer number. The target feature vector is superimposed with the processing result as updated candidate processing data. The next layer of the target layer number is taken as the target layer number of the next iteration according to the layer number of the synthesis module.

[0124] For example, the target layer is the first layer. The processing result corresponding to the first layer of the feature extraction module in the reflection enhancement processing module is obtained. The target feature vector is superimposed with the processing result corresponding to the first layer as the updated candidate processing data. The second layer is used as the target layer for the next iteration.

[0125] The iteration operation continues until the target layer does not have a next layer according to the layer number of the synthesis module. Then the candidate processed data after the iteration operation is used as the enhanced image data.

[0126] By adding skip connections between the image feature processing modules of the synthesis module and the reflection enhancement processing module, the above scheme can better restore image details.

[0127] In some embodiments, the training process of the low-light image enhancement model includes:

[0128] Step 10a: Obtain the initial low-light image enhancement model and training dataset, wherein the training dataset includes low-light image data, target image data, and historical event information corresponding to the low-light image data.

[0129] Step 10b: Input the data in the training dataset into the initial low-light image enhancement model for training until the initial low-light image enhancement model meets the preset training termination condition, then stop training and obtain the low-light image enhancement model.

[0130] In specific implementation, an initial low-light image enhancement model is randomly initialized. This initial low-light image enhancement model includes a reflection component enhancement module, an illumination component enhancement module, and a synthesis module. In this embodiment, the initial low-light image enhancement model can be constructed based on the deep learning framework PyTorch.

[0131] Obtain a training dataset, which includes simulated data and real data. The simulated data consists of the SDSD dataset and event information simulated using V2E, including low-light images, normal-light images, and simulated events. The real data is the DSEC dataset, which includes low-light images and corresponding real events.

[0132] The data from the training dataset is input into the initial low-light image enhancement model for training. Specifically, the low-light images from the training dataset are input into the illumination component enhancement module to extract and enhance the illumination components, resulting in the enhanced illumination components of the original image. The low-light images and events from the training dataset are input into the reflection component enhancement module to extract and enhance the reflection components, resulting in the enhanced reflection components of the original image. The obtained illumination and reflection components are then fused and input into the synthesis module to obtain the enhanced normal light image, which serves as the training result. Training continues until the initial low-light image enhancement model meets a preset training termination condition, at which point the training stops, resulting in the final low-light image enhancement model.

[0133] In some embodiments, the training termination condition includes at least one of the following: all data in the training dataset are input into the initial low-light image enhancement model for training; the number of iterations of the initial low-light image enhancement model reaches a preset number; or the loss function corresponding to the initial low-light image enhancement model converges to a preset convergence threshold.

[0134] For example, the training termination condition is that the number of iterations of the initial low-light image enhancement model reaches a preset number. By comparing the difference between the training results and the normal light image, the model parameters of the initial low-light image enhancement model are adjusted. When the number of iterations reaches the preset number, the training is stopped, and the low-light image enhancement model is obtained.

[0135] For example, rotation and horizontal flipping were used to augment the data, and the network was trained using the AdamW optimizer with a momentum term of (0.9, 0.999). The learning rate was set to 0.001, and a cosine decay strategy was used to reduce it. The network was trained for 200 epochs using an RTX3090. When the preset number of iterations was reached, it indicated that the event-based low-light image enhancement model had achieved good low-light image enhancement capabilities.

[0136] In another example, the training termination condition is that the loss function corresponding to the initial low-light image enhancement model converges to a preset convergence threshold. The loss function of the low-light image enhancement model includes four parts, namely, the temporal consistency loss L. t Detail contrast loss L con Semantic consistency loss L clip Reconstruction loss L error .

[0137] The time consistency loss L t as follows:

[0138] Event Stream E t-Δt~t+Δt The dynamic information between t-Δt and t+Δt is recorded. When Δt→0, it indicates that the motion is very small, reflecting the motion trend. To better maintain the temporal stability of the output frame, a temporal consistency loss L is introduced.t It can be obtained from frame I t The motion trend at time t is estimated and compared with the input. For synthetic data, the input events are aligned with frames, and the temporal consistency loss is:

[0139]

[0140] For real data, the time consistency loss is:

[0141]

[0142] Among them, Y t For the enhanced image, I t The input image at time t, This represents the Frobenius norm, which is a U-Net structure used to extract motion information from an image.

[0143] The detail contrast loss L con as follows:

[0144] Detail information is lost in extremely dark areas, while information in other areas remains relatively intact. To recover complete detail in the enhanced frame, the difference between the input and output frames is increased in the extremely dark areas, while the difference in other areas is reduced. The loss function is defined as follows:

[0145]

[0146] The mask defines the extremely dark area, which is obtained by the region selection module.

[0147] The semantic consistency loss L clip as follows:

[0148] The CLIP model is used to supervise the semantic consistency between the input image and the augmented image, while ensuring that the augmented result conforms to human perception. The loss function is defined as follows:

[0149]

[0150] Where, Φ image and Φ text T represents the image encoder and event encoder in CLIP, respectively, and w represents the weights. n For normal light images, T p This is a low-light image.

[0151] The reconstruction loss L error as follows:

[0152] The reconstruction loss is the sum of squared Euclidean distances between the enhanced image and the corresponding normal, real-light image. The loss function is defined as follows:

[0153]

[0154] Where X t This corresponds to a normal, real light image.

[0155] During training, L is used error and Train the SDSD dataset using L con , and L clip The DSEC dataset is used for training. Therefore, the overall loss function is as follows:

[0156]

[0157] Where α, β, γ, and δ are the trade-off parameters, and the trade-off coefficients are pre-set coefficients.

[0158] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.

[0159] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0160] Based on the same inventive concept, corresponding to any of the above-described embodiments, this disclosure also provides an event-based low-light image enhancement device.

[0161] refer to Figure 2 , Figure 2 An event-based low-light image enhancement apparatus, as described in this embodiment, includes:

[0162] The data acquisition module 201 is configured to acquire initial image data and determine the event information corresponding to the initial image data;

[0163] The data input module 202 is configured to input the initial image data and the event information into a pre-trained low-light image enhancement model, wherein the low-light image enhancement model includes a reflection component enhancement module, an illumination component enhancement module, and a synthesis module;

[0164] The mapping processing module 203 is configured to perform mapping processing on the initial image data and the event information respectively to obtain image feature maps and event feature maps;

[0165] The illumination component enhancement module 204 is configured to input the image feature map into the illumination component enhancement module, and use the illumination component enhancement module to enhance the image feature map to obtain the illumination component.

[0166] The reflection component enhancement module 205 is configured to input the image feature map and the event feature map into the reflection component enhancement module, and use the reflection component enhancement module to enhance the image feature map based on the event feature map to obtain the reflection component;

[0167] The compositing module 206 is configured to input the illumination component and the reflection component into the compositing module, and use the compositing module to perform compositing processing to obtain enhanced image data.

[0168] In some embodiments, the mapping processing module 203 is specifically configured as follows:

[0169] Obtain a preset residual block, input the initial image data and the event information into the residual block, and process the residual block to obtain the initial image feature map and the initial event feature map;

[0170] The initial image data and the initial image feature map are superimposed to obtain the image feature map;

[0171] The event information and the initial event feature map are overlaid to obtain the event feature map.

[0172] In some embodiments, the illumination component enhancement module includes an encoder, and the illumination component enhancement module 204 is specifically configured as follows:

[0173] The image feature map is reduced in dimensionality using an encoder to obtain the target image feature map.

[0174] Component extraction is performed on the feature map of the target image to obtain the initial illumination component;

[0175] The initial illumination component is enhanced to obtain the illumination component.

[0176] In some embodiments, the reflection component enhancement module includes an image feature processing module and an event feature processing module, and the reflection component enhancement module 205 specifically includes:

[0177] The target event feature determination unit is configured to input the event feature map into the event feature processing module, and obtain the target event features through the event feature processing module.

[0178] The reflection component determination unit is configured to input the image feature map and the target event features into the image feature processing module, and use the target event features to enhance the image feature map to obtain the reflection component.

[0179] In some embodiments, the image feature processing module includes a feature extraction module, a downsampling layer, and an image event fusion module, and the reflection component determination unit specifically includes:

[0180] The feature extraction subunit is configured to input the image feature map into the feature extraction module, and use the feature extraction module to extract features from the image feature map to obtain initial image features;

[0181] The amplification processing subunit is configured to input the initial image features into the downsampling layer, and use the downsampling layer to amplify the number of channels of the initial image features to obtain image features;

[0182] The reflection component determination subunit is configured to input the image features and the target event features into the image event fusion module, and use the target event features to enhance the image features to obtain the reflection component.

[0183] In some embodiments, the feature extraction module includes four layers, each layer including a first normalization layer, a first convolutional layer, a channel attention layer, a second normalization layer, and a second convolutional layer. The feature extraction subunit is specifically configured as follows:

[0184] The image feature map is input into the first layer, processed by the first layer to obtain the processing result, and the processing result is input into the next layer. This process is repeated until the image is input into the fourth layer, processed by the fourth layer, and the resulting processing result is used as the initial image feature.

[0185] The specific processing steps at each layer include:

[0186] The image feature map is input into the first normalization layer, processed by the first normalization layer to obtain the first image feature, and the first image feature is sent to the first convolutional layer.

[0187] The first image features are processed using the first convolutional layer to obtain the second image features;

[0188] The second image feature is input into the channel attention layer, and the third image feature is obtained through the channel attention layer.

[0189] The image feature map is superimposed with the third image feature to obtain the fourth image feature, and the fourth image feature is sent to the second normalization layer;

[0190] The fourth image feature is processed using the second normalization layer to obtain the fifth image feature;

[0191] The fifth image feature is input into the second convolutional layer, and the sixth image feature is output after processing by the second convolutional layer.

[0192] The sixth image feature and the fourth image feature are superimposed to obtain the processing result.

[0193] In some embodiments, the image event fusion module includes a region selection module and an attention fusion module, and the reflection component determination subunit is specifically configured as follows:

[0194] The image features are input into the region selection module to determine the image information of multiple image regions corresponding to the image features;

[0195] The image region corresponding to the image information that meets the first preset condition is selected as the target image region, and the target image region is sent to the attention fusion module;

[0196] In the attention fusion module, the target image region is enhanced using target event features to obtain an initial reflection component;

[0197] The initial reflection component is superimposed on the initial image features to obtain the reflection component.

[0198] In some embodiments, the apparatus further includes an alignment processing module, which is specifically configured to:

[0199] The illumination component and the reflection component are input into the alignment module to determine the first channel dimension corresponding to the illumination component and the second channel dimension corresponding to the reflection component.

[0200] Based on the first channel dimension and the second channel dimension, the illumination component and the reflection component are spliced ​​together to obtain the target feature vector.

[0201] In some embodiments, the synthesis processing module 206 is specifically configured as follows:

[0202] The target feature vector is input into the synthesis module, which includes four layers;

[0203] Using the target feature vector as candidate processing data, and taking the first layer of the synthesis module as the target layer number, the candidate processing data is iteratively processed. The specific process of each iteration is as follows:

[0204] Obtain the processing result of the feature extraction module corresponding to the target layer, superimpose the target feature vector with the processing result, and use it as updated candidate processing data. Then, take the next layer of the target layer as the target layer of the next iteration according to the layer of the synthesis module.

[0205] The iteration operation continues until the target layer does not have a next layer according to the layer number of the synthesis module;

[0206] The candidate processed data after iterative operations are used as the enhanced image data.

[0207] In some embodiments, the apparatus further includes a model training module, which is specifically configured to:

[0208] Obtain an initial low-light image enhancement model and a training dataset, wherein the training dataset includes low-light image data, target image data, and historical event information corresponding to the low-light image data;

[0209] The data in the training dataset is input into the initial low-light enhancement model for training until the initial low-light image enhancement model meets the preset training termination condition, at which point training stops and the low-light image enhancement model is obtained.

[0210] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0211] The apparatus of the above embodiments is used to implement the corresponding event-based low-light image enhancement method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0212] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the event-based low-light image enhancement method described in any of the above embodiments.

[0213] Figure 3This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0214] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0215] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0216] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0217] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0218] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0219] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0220] The electronic devices described above are used to implement the corresponding event-based low-light image enhancement methods in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0221] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the event-based low-light image enhancement method as described in any of the above embodiments.

[0222] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0223] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the event-based low-light image enhancement method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0224] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.

[0225] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.

[0226] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0227] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0228] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.

[0229] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuitry) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0230] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0231] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An event-based method for enhancing low-light images, characterized in that, include: Acquire initial image data and determine the event information corresponding to the initial image data; The initial image data and the event information are input into a pre-trained low-light image enhancement model, wherein the low-light image enhancement model includes a reflection component enhancement module, an illumination component enhancement module, and a synthesis module; The initial image data and the event information are respectively mapped to obtain image feature maps and event feature maps; The image feature map is input to the illumination component enhancement module, and the illumination component enhancement module is used to enhance the image feature map to obtain the illumination component. The image feature map and the event feature map are input into the reflection component enhancement module. The reflection component enhancement module enhances the image feature map based on the event feature map to obtain the reflection component. The illumination component and the reflection component are input into the synthesis module, and the synthesis module is used to perform synthesis processing to obtain enhanced image data; The reflection component enhancement module includes an image feature processing module and an event feature processing module. The step of inputting the image feature map and the event feature map into the reflection component enhancement module, and using the reflection component enhancement module to enhance the image feature map based on the event feature map to obtain the reflection component includes: The event feature map is input into the event feature processing module, and the target event features are obtained through the event feature processing module. The image feature map and the target event features are input into the image feature processing module, and the image feature map is enhanced using the target event features to obtain the reflection component. The image feature processing module includes a feature extraction module, a downsampling layer, and an image event fusion module. The step of inputting the image feature map and the target event features into the image feature processing module, and using the target event features to enhance the image feature map to obtain the reflection component includes: The image feature map is input into the feature extraction module, and the feature extraction module is used to extract features from the image feature map to obtain the initial image features; The initial image features are input into a downsampling layer, which amplifies the number of channels of the initial image features to obtain image features. The image features and the target event features are input into the image event fusion module, and the image features are enhanced using the target event features to obtain the reflection component; The image event fusion module includes a region selection module and an attention fusion module. The step of inputting the image features and the target event features into the image event fusion module, and using the target event features to enhance the image features to obtain the reflection component includes: The image features are input into the region selection module to determine the image information of multiple image regions corresponding to the image features; The image region corresponding to the image information that meets the first preset condition is selected as the target image region, and the target image region is sent to the attention fusion module; In the attention fusion module, the target image region is enhanced using target event features to obtain an initial reflection component; The initial reflection component is superimposed on the initial image features to obtain the reflection component.

2. The method according to claim 1, characterized in that, The step of mapping the initial image data and the event information to obtain image feature maps and event feature maps includes: Obtain a preset residual block, input the initial image data and the event information into the residual block, and process the residual block to obtain the initial image feature map and the initial event feature map; The initial image data and the initial image feature map are superimposed to obtain the image feature map; The event information and the initial event feature map are overlaid to obtain the event feature map.

3. The method according to claim 1, characterized in that, The illumination component enhancement module includes an encoder. The step of inputting the image feature map to the illumination component enhancement module and using the illumination component enhancement module to enhance the image feature map to obtain the illumination component includes: The image feature map is reduced in dimensionality using an encoder to obtain the target image feature map. Component extraction is performed on the feature map of the target image to obtain the initial illumination component; The initial illumination component is enhanced to obtain the illumination component.

4. The method according to claim 1, characterized in that, The feature extraction module comprises four layers, each including a first normalization layer, a first convolutional layer, a channel attention layer, a second normalization layer, and a second convolutional layer. The step of inputting the image feature map into the feature extraction module and using the feature extraction module to extract features from the image feature map to obtain initial image features includes: The image feature map is input into the first layer, processed by the first layer to obtain the processing result, and the processing result is input into the next layer. This process is repeated until the image is input into the fourth layer, processed by the fourth layer, and the resulting processing result is used as the initial image feature. The specific processing steps at each layer include: The image feature map is input into the first normalization layer, processed by the first normalization layer to obtain the first image feature, and the first image feature is sent to the first convolutional layer. The first image features are processed using the first convolutional layer to obtain the second image features; The second image feature is input into the channel attention layer, and the third image feature is obtained through the channel attention layer. The image feature map is superimposed with the third image feature to obtain the fourth image feature, and the fourth image feature is sent to the second normalization layer; The fourth image feature is processed using the second normalization layer to obtain the fifth image feature; The fifth image feature is input into the second convolutional layer, and the sixth image feature is output after processing by the second convolutional layer. The sixth image feature and the fourth image feature are superimposed to obtain the processing result.

5. The method according to claim 1, characterized in that, The low-light image enhancement model also includes an alignment module. Before inputting the illumination component and the reflection component into the synthesis module, the method further includes: The illumination component and the reflection component are input into the alignment module to determine the first channel dimension corresponding to the illumination component and the second channel dimension corresponding to the reflection component. Based on the first channel dimension and the second channel dimension, the illumination component and the reflection component are spliced ​​together to obtain the target feature vector; The step of inputting the illumination component and the reflection component into the synthesis module, and using the synthesis module to perform synthesis processing to obtain enhanced image data includes: The target feature vector is input into the synthesis module, which includes four layers; Using the target feature vector as candidate processing data, and taking the first layer of the synthesis module as the target layer number, the candidate processing data is iteratively processed. The specific process of each iteration is as follows: Obtain the processing result of the feature extraction module corresponding to the target layer, superimpose the target feature vector with the processing result, and use it as updated candidate processing data. Then, take the next layer of the target layer as the target layer of the next iteration according to the layer of the synthesis module. The iteration operation continues until the target layer does not have a next layer according to the layer number of the synthesis module, then it exits. The candidate processed data after iterative operations are used as the enhanced image data.

6. The method according to claim 1, characterized in that, The training process of the low-light image enhancement model includes: Obtain an initial low-light image enhancement model and a training dataset, wherein the training dataset includes low-light image data, target image data, and historical event information corresponding to the low-light image data; The data in the training dataset is input into the initial low-light enhancement model for training until the initial low-light image enhancement model meets the preset training termination condition, at which point training stops and the low-light image enhancement model is obtained.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the event-based low-light image enhancement method as described in any one of claims 1 to 6.