Method and device for generating clear enhanced image by fusing image and event under low illumination
By acquiring image and event data under low illumination conditions, cross-modal and cross-scale fusion is performed, and clear and enhanced images are generated using the Swin-CNN module and Transformer block, which solves the problems of traditional camera image blur and low event camera resolution, and achieves high-resolution and clear image generation.
Patent Information
- Application Number
- CN202510833278.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Under low illumination conditions, traditional camera images are blurred, event camera resolution is low and there is strobe phenomenon, and the prior art is difficult to effectively improve image quality.
By acquiring image data and event data under low illumination conditions, extracting multi-resolution image features and event features, performing cross-modal and cross-scale fusion, using Swin-CNN module and Transformer block for feature extraction and fusion, building a loss function to guide model training, and generating clear and enhanced images.
Effectively removes noise from strobe and extremely low illumination conditions, generating high-resolution, clear and enhanced images, solving the problems of traditional camera image blur and event camera low resolution.
Smart Images

Figure CN120339122A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image processing, and particularly relates to a method and device for fusing images and events to generate clear enhanced images under low illumination. Background Art
[0002] Obtaining high-quality images under low illumination conditions has been a long-term challenge. Images taken under low illumination conditions usually suffer from various types of degradation, including poor visibility, noise, and inaccurate colors. Traditional RGB (Red, Green, Blue) cameras, which capture information of the three primary colors to generate true-color images, are difficult to capture sufficient details in low-light environments, resulting in blurred images, low contrast, and loss of details. To overcome these limitations, researchers have explored various image enhancement techniques aimed at improving the visual quality of low-light images. Early methods, such as histogram equalization, improve the overall visual effect by adjusting the contrast of the image. However, these methods only rely on the histogram information of the input image and cannot fully consider the semantic information of the image content, resulting in limited enhancement effects. Deep learning-based low-light image enhancement methods have made significant progress in improving the image quality under low illumination conditions. For example, convolutional neural networks are used for end-to-end mapping to directly convert low-light images into enhanced images. However, due to the low dynamic range and high latency of traditional cameras, these methods are prone to information loss in extremely low light and blurring in moving scenes.
[0003] Event cameras can capture a wider range of brightness by asynchronously detecting the brightness changes of each pixel, thus retaining more details in low-light scenes. Since each pixel independently and immediately responds to changes, event cameras have the characteristic of low latency in generating events and can effectively capture fast-moving targets. Therefore, events can effectively alleviate the limitations faced by current low-light enhancement methods based on image frames. However, under low illumination conditions, low-light enhancement methods that fuse event cameras also face some challenges. On the one hand, the output of event cameras is vulnerable to global stroboscopic interference due to the presence of lights at night, resulting in a large amount of noise. On the other hand, in extremely dark conditions, due to the need for more response time for the same voltage change, trailing noise will be generated. In addition, since the resolution of event cameras is generally lower than that of traditional cameras, it affects the generation of high-resolution images. Summary of the Invention
[0004] The present invention provides a method and device for fusing images and events to generate clear enhanced images under low illumination, so as to solve the problems of blurred images of traditional cameras, low resolution of event cameras, and stroboscopic phenomena in related technologies.
[0005] An embodiment of the first aspect of the present invention provides a method for generating a clear enhanced image by fusing an image and events under low illumination, including the following steps: obtaining image data and event data under low illumination conditions; extracting multi-resolution image features in the image data and multi-resolution event features in the event data, and for the multi-resolution image features and multi-resolution event features, cross-modal selection of new image features and new event features is performed at the same resolution scale; cross-scale fusion is respectively performed on the new image features and new event features of different resolutions to obtain new comprehensive image features and new comprehensive event features, cross-modal fusion is performed on the comprehensive image features and comprehensive event features to obtain fused comprehensive features, and an image with enhanced clarity of the image data is generated based on the fused comprehensive features.
[0006] Optionally, extracting multi-resolution image features in the image data and multi-resolution event features in the event data includes: inputting the image data and event data into a multi-scale encoder, and the multi-scale encoder outputs multi-resolution image features and multi-resolution event features, where the multi-scale encoder includes an image encoder and an event encoder; the image encoder and the event encoder are provided with processing paths at the resolution layer, and the processing path includes a convolutional layer and a convolutional neural network Swin-CNN module enhanced by using a shifted window; the image encoder downsamples the image data to obtain multi-resolution image features, and the event encoder upsamples the event data to obtain multi-resolution event features.
[0007] Optionally, for the multi-resolution image features and multi-resolution event features, cross-modal selection of new image features and new event features at the same resolution scale includes: inputting the multi-resolution image features and multi-resolution event features into a cross-modal feature selection module, and the cross-modal feature selection module outputs new image features and new event features, where the cross-modal feature selection module includes a lower projection layer and an upper projection layer, and the upper projection layer maps the features projected by the lower projection layer back to the corresponding modal space.
[0008] Optionally, the lower projection layer includes a connection layer, a linear layer, a normalization layer, and a fusion layer, where the image features and event features are connected using the connection layer; the linear layer and the normalization layer perform feature dimensionality reduction on the connected features to obtain multi-source features, generate an image prompt and an event prompt according to the multi-source features, and obtain an image source embedding and an event source embedding; generate an image source feature according to the image prompt, the image source embedding, and the image features, and generate an event source feature according to the event prompt, the event source embedding, and the event features; fuse the image source feature and the event source feature to obtain a fused feature, generate a new image feature according to the fused feature, the image features, and the image source embedding, and generate a new event feature according to the fused feature, the event features, and the event source embedding.
[0009] Optionally, cross-scale fusion is performed on the new image features and new event features with different resolutions respectively to obtain new comprehensive image features and new comprehensive event features, including: inputting the new image features with different resolutions and the new event features with different resolutions into a cross-scale feature fusion module, and the cross-scale feature fusion module outputs comprehensive image features and comprehensive event features, where the cross-scale feature fusion module includes a convolutional layer and a Transformer block; at each resolution scale, local information and global information are extracted through the convolutional layer and the Transformer block, features at different resolution scales are aligned through a sampling operation, and features at different resolution scales are fused through the convolutional layer.
[0010] Optionally, cross-modal fusion is performed on the comprehensive image features and comprehensive event features to obtain fused comprehensive features, and an image with enhanced clarity of image data is generated based on the fused comprehensive features, including: inputting the comprehensive image features and comprehensive event features into a cross-modal fusion decoder, and the cross-modal fusion decoder outputs an image with enhanced clarity, where the cross-modal fusion decoder includes a convolutional layer, an activation function, and a Swin-CNN; the convolutional layer and the activation function perform preprocessing on the comprehensive image features and comprehensive event features, and the preprocessed features are input into a first processing channel and a second processing channel; the first processing channel obtains enhanced features through FFT transformation, a Gaussian low-pass filter, an FFC block, and inverse FFT transformation, and the second processing channel is processed through a Swin-CNN block and a convolutional layer, and the features processed by the second processing channel are fused with the enhanced features to obtain fused comprehensive features; the fused comprehensive features are processed through a Swin-CNN and a convolutional layer to obtain an image with enhanced clarity of image data.
[0011] Optionally, a loss function is constructed to guide the training of the model, where the loss function includes a frequency loss, a reconstruction loss, and a total loss. The frequency loss is determined by calculating the norm in the frequency domain, the reconstruction loss is determined by combining the pixel-level error and the structural similarity index, and the total loss is determined by weighting the frequency loss and the reconstruction loss.
[0012] An embodiment of the second aspect of the present invention provides a device for fusing images and events to generate a clearly enhanced image under low illumination, including: an acquisition module for acquiring image data and event data under low illumination conditions; an extraction module for extracting multi-resolution image features in the image data and multi-resolution event features in the event data, and for the multi-resolution image features and multi-resolution event features, cross-modal selection of new image features and new event features is performed at the same resolution scale; a fusion module for performing cross-scale fusion on the new image features and new event features with different resolutions respectively to obtain new comprehensive image features and new comprehensive event features, performing cross-modal fusion on the comprehensive image features and comprehensive event features to obtain fused comprehensive features, and generating an image with enhanced clarity of image data based on the fused comprehensive features.
[0013] In a third aspect embodiment of the present invention, an electronic device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the method for fusing an image and an event under low illumination to generate a clear and enhanced image as described in the above embodiments.
[0014] In a fourth aspect embodiment of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. The program is executed by a processor to implement the method for fusing an image and an event under low illumination to generate a clear and enhanced image as described in the above embodiments.
[0015] Therefore, the present invention has the following beneficial effects: In the embodiments of the present invention, image data and event data are acquired under low illumination conditions; multi-resolution image features in the image data and multi-resolution event features in the event data are extracted. For the multi-resolution image features and multi-resolution event features, new image features and new event features are cross-modally selected at the same resolution scale; the new image features and new event features at different resolutions are respectively cross-scale fused to obtain new comprehensive image features and new comprehensive event features, and the comprehensive image features and comprehensive event features are cross-modally fused to obtain fused comprehensive features. Based on the fused comprehensive features, an image with clear enhancement of the image data is generated. By fusing a high-resolution blurred image and a low-resolution event, the stroboscopic effect and noise under extremely low illumination conditions are effectively removed, and a high-resolution, clear, and enhanced image is generated. Thus, the problems of blurred images of traditional cameras, low resolution of event cameras, and stroboscopic phenomena in the related art are solved.
[0016] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where: Figure 1 is a schematic flowchart of a method for fusing an image and an event under low illumination to generate a clear and enhanced image according to an embodiment of the present invention; Figure 2 is an overall block diagram of a technical solution of a method for fusing an image and an event under low illumination to generate a clear and enhanced image according to an embodiment of the present invention; Figure 3 is a schematic flowchart of a method for fusing an image and an event under low illumination to generate a clear and enhanced image according to an embodiment of the present invention; Figure 4 is a structural diagram of a Swin-CNN multi-scale encoder according to an embodiment of the present invention; Figure 5 The structure diagram of Swin-CNN provided according to an embodiment of the present invention; Figure 6 The structure diagram of SwinT provided according to an embodiment of the present invention; Figure 7 The structure diagram of the feature dimension reduction layer provided according to an embodiment of the present invention; Figure 8 The structure diagram of the fusion layer provided according to an embodiment of the present invention; Figure 9 The structure diagram of the cross-scale feature fusion module based on the sparse transformer block provided according to an embodiment of the present invention; Figure 10 The block diagram of SparSTB and SparCTB provided according to an embodiment of the present invention; Figure 11 The structure diagram of the cross-modal fusion decoder based on frequency domain filtering provided according to an embodiment of the present invention; Figure 12 The block diagram of the device for generating a clear enhanced image by fusing an image and events under low illumination according to an embodiment of the present invention; Figure 13 The structure diagram of the electronic device according to an embodiment of the present invention. Detailed implementation manners
[0018] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0019] The method and device for generating a clear enhanced image by fusing an image and events under low illumination according to the embodiments of the present invention will be described below with reference to the accompanying drawings. Aiming at the problems of blurred images of traditional cameras, low resolution and stroboscopic phenomena in event cameras in the related art mentioned in the above background art, the present invention provides a method for generating a clear enhanced image by fusing an image and events under low illumination. In this method, image data and event data are acquired under low illumination conditions; multi-resolution image features in the image data and multi-resolution event features in the event data are extracted. For the multi-resolution image features and multi-resolution event features, new image features and new event features are cross-modally selected at the same resolution scale; the new image features and new event features at different resolutions are respectively cross-scale fused to obtain new comprehensive image features and new comprehensive event features, the comprehensive image features and comprehensive event features are cross-modally fused to obtain fused comprehensive features, and an image with enhanced clarity of the image data is generated based on the fused comprehensive features. By fusing a high-resolution blurred image and a low-resolution event, a clear image with high resolution and normal brightness is generated. Thus, the problems of blurred images of traditional cameras, low resolution and stroboscopic phenomena in event cameras in the related art are solved.
[0020] Specifically, Figure 1 FIG. is a schematic flowchart of a method for generating a clear enhanced image by fusing an image and events under low illumination according to an embodiment of the present invention.
[0021] As Figure 1 shown, the method for generating a clear enhanced image by fusing an image and events under low illumination includes the following steps: In step S101, image data and event data are acquired under low illumination conditions.
[0022] Among them, the low illumination condition refers to an environmental condition where the ambient light is very weak; the image data refers to the information captured by a traditional RGB camera, including visual information such as color and brightness, etc.; the event data refers to the information captured by an event camera. An event camera is a bio-inspired visual sensor that works by asynchronously detecting the brightness changes of each pixel.
[0023] In step S102, multi-resolution image features in the image data and multi-resolution event features in the event data are extracted. For the multi-resolution image features and multi-resolution event features, new image features and new event features are cross-modally selected at the same resolution scale.
[0024] Among them, the multi-resolution image features refer to the feature information at different scales extracted from the image data, including high, medium, and low resolution image features; the multi-resolution event features refer to the features extracted at different time resolutions according to the information output by the event camera, including high, medium, and low resolution event features.
[0025] It can be understood that in the embodiments of the present invention, multi-resolution image features and multi-resolution event features are extracted from the image data and event data obtained under low illumination conditions, and new image features and new event features are cross-modally selected at the same resolution scale. The specific selection method will be described in detail below and will not be elaborated here.
[0026] In the embodiments of the present invention, extracting multi-resolution image features from image data and multi-resolution event features from event data includes: inputting the image data and event data into a multi-scale encoder, and the multi-scale encoder outputs multi-resolution image features and multi-resolution event features. Among them, the multi-scale encoder includes an image encoder and an event encoder; the image encoder and the event encoder are provided with processing paths at the resolution layer, and the processing paths include a convolutional layer and a convolutional neural network using a shifted window for enhancement, namely, the Swin-CNN module; the image encoder downsamples the image data to obtain multi-resolution image features, and the event encoder upsamples the event data to obtain multi-resolution event features.
[0027] Among them, Swin-CNN is a convolutional neural network using a shifted window for enhancement; the Swin-CNN module includes a convolutional layer and a SwinT (Shifted Window Multi-head Self-Attention) module, and the SwinT module further enhances the feature representation ability through components such as a multi-layer perceptron, a normalization layer, window multi-head self-attention, and shifted window multi-head self-attention.
[0028] It can be understood that the embodiments of the present invention use a multi-scale encoder to obtain multi-resolution image features and multi-resolution event features based on image data and event data. Specifically, the multi-scale encoder includes an image encoder and an event encoder. The image encoder receives a high-resolution traditional image as input and generates multi-resolution hierarchical image features through a downsampling operation; the event encoder receives low-resolution event data input and obtains multi-resolution event features through an upsampling operation. The downsampled image and the upsampled event can achieve cross-scale alignment of high-resolution images and low-resolution events, while ensuring correct alignment of features at different scales, avoiding scale conversion, and making subsequent feature selection more straightforward.
[0029] In the embodiments of the present invention, for multi-resolution image features and multi-resolution event features, cross-modal selection of new image features and new event features at the same resolution scale includes: inputting the multi-resolution image features and multi-resolution event features into a cross-modal feature selection module, and the cross-modal feature selection module outputs new image features and new event features. Among them, the cross-modal feature selection module includes a down-projection layer and an up-projection layer, and the up-projection layer maps the features projected by the down-projection layer back to the corresponding modal space.
[0030] Among them, the down-projection layer and the up-projection layer will be described in detail below and will not be elaborated here.
[0031] It can be understood that in the embodiments of the present invention, the cross-modal feature selection module outputs new image features and new event features. In this cross-modal feature selection module, there are a down-projection layer and an up-projection layer. The up-projection layer maps the features projected by the down-projection layer back to the corresponding modal space, thereby forming improved new image features and new event features. Through cross-modal feature selection, the most beneficial features can be extracted and fused from the two modalities to generate more accurate and detailed new image features and new event features.
[0032] In the embodiments of the present invention, the down-projection layer includes a connection layer, a linear layer, a normalization layer, and a fusion layer. Among them, the image features and event features are connected using the connection layer; the linear layer and the normalization layer perform feature dimensionality reduction on the connected features to obtain multi-source features, generate image cues and event cues based on the multi-source features, and obtain image source embeddings and event source embeddings; generate image source features based on the image cues, image source embeddings, and image features, and generate event source features based on the event cues, event source embeddings, and event features; fuse the image source features and event source features to obtain fusion features, generate new image features based on the fusion features, image features, and image source embeddings, and generate new event features based on the fusion features, event features, and event source embeddings.
[0033] Among them, the image source embedding is the basic representation extracted from the input high-resolution blurred image, which contains information such as the basic structure, color, and texture of the image; the event source embedding is the basic representation extracted from the input low-resolution event data, and the event data reflects the time series information of pixel brightness changes.
[0034] It can be understood that in the embodiment of the present invention, in the lower projection layer of the cross-modal feature selection module, the connection layer is used to connect the image feature and the event feature to obtain the connected feature, and then the linear layer and the normalization layer are used to reduce the dimension of the connected feature to obtain the multi-source feature. Based on the multi-source feature, image prompts and event prompts are generated, and these prompts are used together with the image source embedding and the event source embedding to generate the refined image source feature and event source feature. The image source feature and the event source feature are fused to obtain the fused feature, and finally, new image features are generated according to the fused feature, the image feature, and the image source embedding, and new event features are generated according to the fused feature, the event feature, and the event source embedding.
[0035] In step S103, the new image features and new event features with different resolutions are respectively subjected to cross-scale fusion to obtain new comprehensive image features and new comprehensive event features. The comprehensive image features and the comprehensive event features are subjected to cross-modal fusion to obtain the fused comprehensive feature, and an image with enhanced clarity of the image data is generated based on the fused comprehensive feature.
[0036] Among them, cross-scale fusion refers to integrating features at different resolution levels to obtain a more comprehensive information representation; cross-modal fusion refers to fusing data from different modalities, such as images and events, with the aim of combining their respective advantages to generate a more accurate and rich feature representation.
[0037] It can be understood that in the embodiment of the present invention, by respectively performing cross-scale fusion on the new image features and new event features with different resolutions, optimized new comprehensive image features and new comprehensive event features are obtained, aiming to integrate multi-level information, thereby capturing more detailed local details and global structures; the comprehensive image features and the comprehensive event features are subjected to cross-modal fusion, and the fused comprehensive feature is further refined by using their respective advantages. Finally, an image with enhanced clarity of the image data, that is, an image with higher resolution and better visual quality, is generated based on the fused comprehensive feature.
[0038] In the embodiment of the present invention, the new image features and new event features with different resolutions are respectively subjected to cross-scale fusion to obtain new comprehensive image features and new comprehensive event features, including: inputting the new image features with different resolutions and the new event features with different resolutions into the cross-scale feature fusion module, and the cross-scale feature fusion module outputs the comprehensive image features and the comprehensive event features. Among them, the cross-scale feature fusion module includes a convolutional layer and a Transformer block; at each resolution scale, local information and global information are extracted through the convolutional layer and the Transformer block, the features at different resolution scales are aligned through sampling operations, and the features at different resolution scales are fused through the convolutional layer.
[0039] Among them, the Transformer block is a sparse Transformer block, including a spatial-window-based sparse Transformer block (SparSTB) and a channel-window-based sparse Transformer block (SparCTB). Specifically, the spatial-window-based Transformer block (STB) is a self-attention mechanism that reduces computational complexity and improves processing efficiency by dividing the input image into multiple local windows and performing self-attention calculations within each window. The channel-based Transformer block (CTB) is an important variant of the Transformer in the field of computer vision. In particular, it improves computational efficiency by focusing on the attention mechanism in the channel dimension while maintaining a good ability to model the spatial structure of the image. SparSTB and SparCTB are the spatial-window-based sparse Transformer block and the channel-window-based sparse Transformer block designed based on STB and CTB respectively.
[0040] It can be understood that in the embodiments of the present invention, new image features and new event features with different resolutions are input into the cross-scale feature fusion module. In the cross-scale feature fusion module, features at each resolution level are processed through convolutional layers and Transformer blocks to extract local and global information. Sampling operations are adopted to ensure that features at different resolution scales can be correctly aligned. The features at different resolution scales are fused through convolutional layers. Finally, the cross-scale feature fusion module outputs comprehensive image features and comprehensive event features, providing a basis for subsequent cross-modal fusion steps.
[0041] In an embodiment of the present invention, cross-modal fusion is performed on the comprehensive image features and comprehensive event features to obtain fused comprehensive features, and an image with enhanced clarity of image data is generated based on the fused comprehensive features, including: inputting the comprehensive image features and comprehensive event features into a cross-modal fusion decoder, and the cross-modal fusion decoder outputs an image with enhanced clarity, where the cross-modal fusion decoder includes a convolutional layer, an activation function, and a Swin-CNN; the convolutional layer and the activation function perform preprocessing on the comprehensive image features and comprehensive event features, and the preprocessed features are respectively input into a first processing channel and a second processing channel; the first processing channel obtains enhanced features through FFT (Fast Fourier Transform), a Gaussian low-pass filter, an FFC block (Fast Fourier Convolution Block), and an inverse FFT transform, and the second processing channel is processed through a Swin-CNN block and a convolutional layer, and the features processed by the second processing channel are fused with the enhanced features to obtain fused comprehensive features; the fused comprehensive features are processed through a Swin-CNN and a convolutional layer to obtain an image with enhanced clarity of image data.
[0042] Among them, the activation function can be a LeakyReLU (Leaky Rectified Linear Unit) function; the FFT transform is used to convert a time-domain signal into a frequency-domain signal for frequency-domain filtering; the inverse FFT transform is the inverse process of the fast Fourier transform, which converts the processed frequency-domain signal back to the time domain or spatial domain; the Gaussian low-pass filter is used to filter out high-frequency noise and retain low-frequency information, which helps to smooth the image features; the FFC block is a convolution operation implemented using the fast Fourier transform, which can perform efficient feature selection in the frequency domain.
[0043] It can be understood that in the embodiment of the present invention, the comprehensive image features and comprehensive event features are input into the cross-modal fusion decoder. Inside the cross-modal fusion decoder, the comprehensive image features and comprehensive event features are preprocessed by the convolutional layer and the activation function, and then input into two processing channels: the first processing channel processes the features through the fast Fourier transform FFT, a Gaussian low-pass filter, an FFC block, and an inverse FFT transform to remove noise and highlight low-frequency information, thereby obtaining enhanced features; the second processing channel processes the features through a Swin-CNN block and a convolutional layer, combines the features obtained from the second processing channel with the enhanced features of the first processing channel to form fused comprehensive features, and these fused comprehensive features are processed again through a Swin-CNN and a convolutional layer to finally generate an image data with enhanced clarity, effectively removing noise, retaining important details, and effectively combining the advantages of frequency-domain and spatial-domain processing, achieving high-quality low-light image enhancement.
[0044] In an embodiment of the present invention, a loss function is constructed to guide the training of the model. The loss function includes a frequency loss, a reconstruction loss, and a total loss. The frequency loss is determined by calculating the norm in the frequency domain, the reconstruction loss is determined by combining the pixel-level error and the structural similarity index, and the total loss is determined by weighting the frequency loss and the reconstruction loss.
[0045] Specifically, the loss function is , where is the final network output, is the corresponding ground truth of the front view, =0.1 , is the frequency loss, , denotes the Fourier transform; is the reconstruction loss, which consists of two parts: pixel error and structural consistency constraint. , where is the enhancement result, is the ground truth, is the structural similarity index between the two, and the norm is used to calculate the absolute pixel error between the enhancement result and the ground truth . The structural similarity index is also used to compare the similarity of the images.
[0046] It can be understood that in the embodiment of the present invention, the frequency loss is first determined by calculating the norm in the frequency domain to reduce the high-frequency noise in the output image, the reconstruction loss is obtained by combining the pixel-level error and the structural similarity index, and finally, the total loss is calculated by assigning appropriate weights to the frequency loss and the reconstruction loss. The total loss is used to supervise the model, so that the output image is highly similar to the real image in both the spatial and frequency domains, thereby effectively guiding the model training and improving the quality of image enhancement under low illumination conditions.
[0047] A method for generating a clear and enhanced image by fusing an image and events under low illumination according to an embodiment of the present invention includes obtaining image data and event data under low illumination conditions; extracting multi-resolution image features from the image data and multi-resolution event features from the event data. For the multi-resolution image features and multi-resolution event features, new image features and new event features are cross-modally selected at the same resolution scale; the new image features and new event features at different resolutions are respectively cross-scale fused to obtain new comprehensive image features and new comprehensive event features, and the comprehensive image features and comprehensive event features are cross-modally fused to obtain fused comprehensive features. An image with enhanced clarity of the image data is generated based on the fused comprehensive features. By fusing a high-resolution blurred image and low-resolution events, stroboscopic and noise under extremely low illumination conditions are effectively removed, and a high-resolution, clear, and enhanced image is generated.
[0048] The method for generating a clear and enhanced image by fusing an image and events under low illumination is further described below through a specific embodiment.
[0049] Figure 2 The overall block diagram of the technical solution of a method for generating a clear and enhanced image by fusing an image and events under low illumination provided in this embodiment is Figure 3 The flowchart of a method for generating a clear and enhanced image by fusing an image and events under low illumination provided in this embodiment is as Figure 3 shown. The specific implementation scheme of this embodiment is as follows: Step S201: Build a Swin-CNN multi-scale encoder The Swin-CNN multi-scale image encoder based on downsampling and the Swin-CNN multi-scale event encoder based on upsampling are as Figure 4 shown.
[0050] First, construct a low-resolution event representation. The low-resolution event includes two parts, namely the time-dependent representation frame and the full-time event representation frame . The time-dependent representation frame contains event information within centered on , and the full-time event representation frame contains event information within the entire exposure time .
[0051]
[0052] Where represents pixel-by-pixel superposition.
[0053] Low-resolution events combine event information with different time resolutions. They can not only obtain clear events with low latency, which can be used to supplement the information loss of traditional cameras and remove the blur of traditional cameras, but also acquire the overall events during the entire exposure time, providing overall information for subsequent reduction of stroboscopic effects and the influence of noise in event cameras under extremely low illumination conditions.
[0054] Input a high-resolution image and perform downsampling on the image. The images and events at each scale are processed through their respective convolutional layers and Swin-CNN to generate high-, medium-, and low-resolution image features. Similarly, input low-resolution events and perform upsampling on the events to generate high-, medium-, and low-resolution event features.
[0055] The downsampled images and upsampled events can achieve cross-scale alignment of high-resolution images and low-resolution events. At the same time, it ensures the correct alignment of features at different scales, avoids scale conversion, and makes subsequent feature selection more straightforward. Features at different scales have different focuses. Higher-resolution features are beneficial for extracting shallow features, and lower-resolution features are beneficial for extracting deep features. Multi-scale feature extraction can comprehensively extract the information of images and events.
[0056] The Swin-CNN module is the core of this process. The module, as Figure 5 shown, contains a convolutional layer and a SwinT module to ensure the effective extraction of local and global features. The SwinT module is as Figure 6 shown. The SwinT module further enhances the feature representation ability of the model through components such as MLP (Multilayer Perceptron), LayerNorm, W-MSA (Window-based Multi-head Self-Attention), and SW-MSA (Shifted Window Multi-head Self-Attention).
[0057] Step S202: Build a prompt-driven cross-modal feature selection module Under low illumination conditions, images provide color and texture information, while events provide high-dynamic detail information. Through cross-modal feature selection, the most beneficial features can be extracted and fused from the two modalities to generate a more accurate and detailed representation. Operating at the same scale can ensure that features from different modalities are aligned in spatial resolution. Under challenging conditions such as low illumination, the system can more accurately identify and locate objects in the scene.
[0058] Feature selection is divided into four steps: multi-source feature generation, prompt generation, prompt-driven fusion, and mutual information regularization.
[0059] First, generate multi-source features. Input image features and event features , and use for feature dimensionality reduction to obtain the processed multi-source features
[0060]
[0061] where represents the concatenation of features, including a linear layer and a normalization layer, as Figure 7 shown.
[0062] Then, generate prompts,
[0063] where represents global average pooling, represents an activation function, represents a cross-modal adapter, represents the gating network corresponding to the adapter, and the prompt consists of an image prompt and an event prompt .
[0064] The gating network is defined as
[0065] where and represent activation functions, represents keeping only the first values ( ), and are learnable parameters, is the input multi-source feature.
[0066] The cross-modal adapter is defined as
[0067]
[0068] The cross-modal adapter can achieve the interaction between multi-modal information. By sharing partial weights, multiple modalities are unified into a single framework. Its main goal is to enhance the interaction between different modalities while reducing the number of additional tunable parameters. represents the image feature, represents the event feature, represents the scaling factor. The unified lower projection layer is a shared layer used to project input features into a bottleneck layer; a non-linear activation function is used to introduce non-linearity and enhance the expressive power of the model; modality-specific up-projection layers These layers are specific to each modality and are used to map the down-projected features back to their respective modality spaces. In the multi-modal case, an additional cross-modal up-projection layer is introduced to effectively handle the mixed information.
[0069] Image prompt and event prompt are respectively dot-multiplied with the source feature image feature and event feature by learnable parameters source embeddings independent of the input, namely source embeddings and , to generate refined source features and , and finally pass through a fusion layer to obtain a fused feature . Using prompt-driven dynamic adjustment of the weights of different modalities to achieve feature selection, so as to improve the adaptability of the network to complex scenarios such as stroboscopic and extremely low illuminance.
[0070]
[0071]
[0072]
[0073] Among them, the fusion layer uses an addition operation to fuse the refined multi-source features, and then passes through a set of convolutional layers. The structure of the fusion layer is as Figure 8 shown.
[0074] The features input to the next module are processed as follows,
[0075]
[0076] Among them represents the image feature and represents the event feature, represents the fused feature, is a learnable parameter initialized to 0.5.
[0077] To ensure that the model dynamically retains complementary information while discarding redundant information of multi-source features, for the image prompt and event prompt Apply a regularization constraint to enforce a balanced information distribution among the prompts. Define Mutual Information Regularization (MIR) as follows:
[0078] Step S203: Build a cross-scale feature fusion module based on sparse transformer blocks Input high-, medium-, and low-resolution image features and event features. The features at each resolution are processed through a 3x3 convolutional layer (Conv3x3) and sparse transformer blocks. The features are concatenated at different levels and further fused through a 1x1 convolutional layer (Conv1x1). The low-resolution and medium-resolution features are aligned with the high-resolution features through upsampling operations. Finally, the fused features pass through another 3x3 convolutional layer to generate the final feature representation. This cross-scale processing and fusion process aims to enhance the feature representation by combining information at different scales, providing cross-scale fused image features and event features for subsequent cross-modal feature fusion. The module structure is as Figure 9 shown.
[0079] The Spatial Window-based Transformer Block (STB) is a self-attention mechanism that reduces computational complexity and improves processing efficiency by dividing the input image into multiple local windows and performing self-attention calculations within each window. The Channel-based Transformer Block (CTB) is an important variant of Transformer in the field of computer vision. In particular, it improves computational efficiency by focusing on the attention mechanism in the channel dimension while maintaining a good modeling ability for the spatial structure of the image.
[0080] In this embodiment, a Sparse Block is designed based on STB and CTB, and a Spatial Window-based Sparse Transformer Block (SparSTB) and a Channel Window-based Sparse Transformer Block (SparCTB) are designed based on the Sparse Block, as Figure 10 shown. The Sparse Block is used to convert the dimension size of the output from the previous stage into new tokens with a new embedding dimension size. The conversion process is to obtain a new embedding representation through a convolutional layer and convert it into latent tokens through a linear layer.
[0081] In this embodiment, the convolutional kernel and stride size used in the convolutional layer are 3 and 1, respectively. At the same time, the input features in the linear layer are in the third stage, and the output features are 49 to generate new token representations. The output obtained from the SparseBlock is where is the number of new tags required, is the number of new embedding dimensions.
[0082] Step S204: Cross-modal fusion decoder based on frequency-domain filtering In this embodiment, a cross-modal fusion decoder based on frequency-domain filtering is designed, with event features and image features as inputs, processed through a 1x1 convolutional layer and a LeakyReLU activation function, and then enter the low-frequency channel and the high-frequency channel. The low-frequency channel includes a convolutional layer, an FFT transform, a Gaussian low-pass filter, and an FFC block for frequency-domain filtering, and then through an inverse FFT transform and a per-pixel spatial dynamic filter, processed through a frequency adaptive fusion module and a further convolutional layer to generate low-frequency features. The high-frequency features are processed through a convolutional layer, a Swin-CNN module, and a further convolutional layer, and fused with the low-frequency features. Finally, the fused features pass through a Swin-CNN and a convolutional layer to generate a high-resolution clear enhanced image. This multi-stage processing method aims to effectively remove noise, retain important details, and generate a clearer and more accurate image representation. The module structure is as Figure 11 shown.
[0083] Under low-light conditions, a low signal-to-noise ratio will cause significant noise in the image, making it extremely difficult to recover the global structural information. In addition, in stroboscopic scenes and extremely low-light scenes, events also tend to generate a large amount of noise. Therefore, a feature fusion enhancement method based on frequency-domain filtering is designed, which can effectively reduce this high-frequency noise and accurately recover the main structural information of the scene.
[0084] A Gaussian low-pass filter is used to extract low-frequency information. The Gaussian filter is defined as follows:
[0085] where, and respectively represent the center points of the axis and the axis in the frequency domain, represents the standard deviation, , are the coordinates in the frequency domain. After low-pass filtering, an additional frequency-domain selection is performed through a fast Fourier convolution (FFC) block. Finally, an inverse fast Fourier transform (Inverse-FFT) is applied for inverse transformation to generate the frequency-filtered features. The frequency-filtered features are combined with the original features through a residual connection. These features emphasized by low-frequency information usually correspond to the spatially varying main structural information. To better enhance the main structure of the scene, a per-pixel spatial dynamic filter is applied to the features.
[0086] Since the fusion process pays more attention to the complementary information of the source images, the residuals of the features of the two source images are first calculated. Subsequently, the residuals are input into the Swin-CNN module to extract semantic features. Then, a single-layer convolution is used to map the multi-channel semantic features into two-channel features. Finally, a SoftMax operation is applied to normalize the features to generate two masks. These masks are then multiplied element-wise with the features of the source images to obtain the fused features.
[0087] After obtaining the low-frequency fused features and high-frequency fused features, they will respectively pass through a convolutional layer and a Sigmoid to remove irrelevant spatial information, and then be added and decoded through a Swin-CNN and a convolutional layer to obtain a high-resolution clear enhanced image.
[0088] Step S205: Construct a loss function The network is supervised by a loss function, and the input of this loss function includes the final network output and the corresponding ground truth of the main view . The loss function designed in the present invention includes two parts: frequency loss and reconstruction loss .
[0089]
[0090] Among them, ; the frequency loss function , specifically for blurred images, suppressing high-frequency noise, can be expressed as
[0091] denotes the Fourier transform.
[0092] Among them, the reconstruction loss function consists of two parts: pixel error and structural consistency constraint. The reconstruction loss can be expressed as
[0093] Adopt norm to calculate the absolute pixel error between the enhanced result and the ground truth , and also use the structural similarity index to compare the similarity of images.
[0094] Secondly, a device for clear enhanced images under low illumination according to an embodiment of the present invention will be described with reference to the accompanying drawings.
[0095] Figure 12 is a schematic block diagram of a device for generating clear enhanced images by fusing images and events under low illumination provided by an embodiment of the present invention.
[0096] As Figure 12 shown, the apparatus 10 for generating a clear enhanced image by fusing an image and events under low illumination includes: an acquisition module 301, an extraction module 302, and a fusion module 303.
[0097] Among them, the acquisition module 301 is used to acquire image data and event data under low illumination conditions; the extraction module 302 is used to extract multi-resolution image features in the image data and multi-resolution event features in the event data. For the multi-resolution image features and multi-resolution event features, new image features and new event features are cross-modally selected at the same resolution scale; the fusion module 303 is used to perform cross-scale fusion on the new image features and new event features of different resolutions respectively to obtain new comprehensive image features and new comprehensive event features, perform cross-modal fusion on the comprehensive image features and comprehensive event features to obtain fused comprehensive features, and generate an image with enhanced clarity of the image data based on the fused comprehensive features.
[0098] In the embodiment of the present invention, when extracting multi-resolution image features in the image data and multi-resolution event features in the event data, the extraction module 302 is further configured to: input the image data and the event data into a multi-scale encoder, and the multi-scale encoder outputs multi-resolution image features and multi-resolution event features. Among them, the multi-scale encoder includes an image encoder and an event encoder; the image encoder and the event encoder are provided with processing paths at the resolution layer, and the processing paths include a convolutional layer and a convolutional neural network Swin-CNN module enhanced by using a shifted window; the image encoder downsamples the image data to obtain multi-resolution image features, and the event encoder upsamples the event data to obtain multi-resolution event features.
[0099] In the embodiment of the present invention, for the multi-resolution image features and multi-resolution event features, when cross-modally selecting new image features and new event features at the same resolution scale, the extraction module 302 is further configured to: input the multi-resolution image features and multi-resolution event features into a cross-modal feature selection module, and the cross-modal feature selection module outputs new image features and new event features. Among them, the cross-modal feature selection module includes a lower projection layer and an upper projection layer, and the upper projection layer maps the features projected by the lower projection layer back to the corresponding modal space.
[0100] In an embodiment of the present invention, the lower projection layer includes a connection layer, a linear layer, a normalization layer, and a fusion layer. Among them, the image feature and the event feature are connected using the connection layer; the linear layer and the normalization layer perform feature dimensionality reduction on the connected features to obtain multi-source features, generate an image prompt and an event prompt according to the multi-source features, and obtain an image source embedding and an event source embedding; generate an image source feature according to the image prompt, the image source embedding, and the image feature, and generate an event source feature according to the event prompt, the event source embedding, and the event feature; fuse the image source feature and the event source feature to obtain a fusion feature, generate a new image feature according to the fusion feature, the image feature, and the image source embedding, and generate a new event feature according to the fusion feature, the event feature, and the event source embedding.
[0101] In an embodiment of the present invention, cross-scale fusion is performed on the new image features and new event features with different resolutions respectively to obtain new comprehensive image features and new comprehensive event features. The fusion module 303 is further configured to: input the new image features with different resolutions and the new event features with different resolutions into a cross-scale feature fusion module, and the cross-scale feature fusion module outputs comprehensive image features and comprehensive event features. Among them, the cross-scale feature fusion module includes a convolutional layer and a Transformer block; at each resolution scale, local information and global information are extracted through the convolutional layer and the Transformer block, the features at different resolution scales are aligned through a sampling operation, and the features at different resolution scales are fused through the convolutional layer.
[0102] In an embodiment of the present invention, cross-modal fusion is performed on the comprehensive image features and comprehensive event features to obtain a fused comprehensive feature, and an image with enhanced clarity of image data is generated based on the fused comprehensive feature. The fusion module 303 is further configured to: input the comprehensive image features and comprehensive event features into a cross-modal fusion decoder, and the cross-modal fusion decoder outputs an image with enhanced clarity. Among them, the cross-modal fusion decoder includes a convolutional layer, an activation function, and a Swin-CNN; the convolutional layer and the activation function perform preprocessing on the comprehensive image features and comprehensive event features, and the preprocessed features are input into a first processing channel and a second processing channel; the first processing channel obtains an enhanced feature through FFT transformation, a Gaussian low-pass filter, an FFC block, and inverse FFT transformation, and the second processing channel is processed through a Swin-CNN block and a convolutional layer, and the features processed by the second processing channel are fused with the enhanced feature to obtain a fused comprehensive feature; the fused comprehensive feature is processed through a Swin-CNN and a convolutional layer to obtain an image with enhanced clarity of image data.
[0103] In an embodiment of the present invention, a construction module is further included. The construction module is further configured to construct a loss function, and use the loss function to guide the training of the network. The loss function includes a frequency loss, a reconstruction loss, and a total loss. The frequency loss is determined by calculating a norm in the frequency domain, the reconstruction loss is determined by combining pixel-level errors and a structural similarity index, and the total loss is determined by weighting the frequency loss and the reconstruction loss.
[0104] It should be noted that the foregoing explanations of the method embodiments for fusing images and events to generate clear enhanced images under low illumination also apply to the device for fusing images and events to generate clear enhanced images under low illumination in this embodiment, and will not be elaborated here.
[0105] The device for fusing images and events to generate clear enhanced images according to an embodiment of the present invention obtains image data and event data under low illumination conditions; extracts multi-resolution image features in the image data and multi-resolution event features in the event data. For the multi-resolution image features and multi-resolution event features, new image features and new event features are cross-modally selected at the same resolution scale; the new image features and new event features at different resolutions are respectively cross-scale fused to obtain new comprehensive image features and new comprehensive event features, the comprehensive image features and comprehensive event features are cross-modally fused to obtain fused comprehensive features, and an image with clear enhancement of the image data is generated based on the fused comprehensive features. By fusing high-resolution blurred images and low-resolution events, stroboscopic and noise under extremely low illumination conditions are effectively removed, and high-resolution, clear, and enhanced images are generated.
[0106] Figure 13 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device may include: A memory 401, a processor 402, and a computer program stored on the memory 401 and executable on the processor 402.
[0107] When the processor 402 executes the program, it implements the method for fusing images and events to generate clear enhanced images under low illumination provided in the foregoing embodiment.
[0108] Further, the electronic device further includes: A communication interface 403 for communication between the memory 401 and the processor 402.
[0109] The memory 401 is used to store a computer program executable on the processor 402.
[0110] The memory 401 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.
[0111] If the memory 401, the processor 402, and the communication interface 403 are implemented independently, the communication interface 403, the memory 401, and the processor 402 can be interconnected via a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, Figure 13 it is represented by only a thick line in the figure, but it does not mean that there is only one bus or one type of bus.
[0112] Optionally, in specific implementation, if the memory 401, the processor 402, and the communication interface 403 are integrated on a single chip, the memory 401, the processor 402, and the communication interface 403 can communicate with each other via an internal interface.
[0113] The processor 402 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0114] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method for fusing an image and an event under low illumination to generate a clear enhanced image as described above is implemented.
[0115] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0116] In addition, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0117] Any process or method description shown in a flowchart or described otherwise herein can be understood to represent a module, segment, or portion of code including one or N executable instructions for implementing a customized logical function or process. And the scope of the preferred embodiments of the present invention includes additional implementations where functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0118] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays, field programmable gate arrays, etc.
[0119] Those of ordinary skill in the art of the present technology can understand that all or part of the steps carried by the methods for implementing the above embodiments can be completed by instructing relevant hardware through a program. The above program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0120] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for generating a clear enhanced image by fusing an image and events under low illumination, characterized in that, Including the following steps: Obtain image data and event data under low-light conditions; Extract multi-resolution image features in the image data and multi-resolution event features in the event data. For the multi-resolution image features and multi-resolution event features, cross-modal select new image features and new event features at the same resolution scale; Perform cross-scale fusion on the new image features and new event features of different resolutions respectively to obtain new comprehensive image features and new comprehensive event features, perform cross-modal fusion on the comprehensive image features and the comprehensive event features to obtain fused comprehensive features, and generate an image with enhanced clarity of the image data based on the fused comprehensive features.
2. The method for generating a clear enhanced image by fusing an image and events under low illumination according to claim 1, wherein, The extracting multi-resolution image features in the image data and multi-resolution event features in the event data includes: Input the image data and event data into a multi-scale encoder, and the multi-scale encoder outputs multi-resolution image features and multi-resolution event features, where the multi-scale encoder includes an image encoder and an event encoder; The image encoder and the event encoder are provided with processing paths at the resolution layer, and the processing paths include a convolutional layer and a convolutional neural network Swin-CNN module enhanced by using a shifted window; The image encoder downsamples the image data to obtain multi-resolution image features, and the event encoder upsamples the event data to obtain multi-resolution event features.
3. The method for generating a clear enhanced image by fusing an image and events under low illumination according to claim 1, wherein The cross-modal selecting new image features and new event features at the same resolution scale for the multi-resolution image features and multi-resolution event features includes: Input the multi-resolution image features and multi-resolution event features into a cross-modal feature selection module, and the cross-modal feature selection module outputs new image features and new event features, where the cross-modal feature selection module includes a lower projection layer and an upper projection layer, and the upper projection layer maps the features projected by the lower projection layer back to the corresponding modal space.
4. The method for generating a clear enhanced image by fusing an image and events under low illumination according to claim 3, wherein The lower projection layer includes a connection layer, a linear layer, a normalization layer, and a fusion layer, where The image features and the event features are connected using the connection layer; The linear layer and the normalization layer perform feature dimensionality reduction on the connected features to obtain multi-source features, generate an image prompt and an event prompt according to the multi-source features, and obtain an image source embedding and an event source embedding; Generate an image source feature according to the image prompt, the image source embedding, and the image features, and generate an event source feature according to the event prompt, the event source embedding, and the event features; Fuse the image source feature and the event source feature to obtain a fused feature, generate a new image feature according to the fused feature, the image features, and the image source embedding, and generate a new event feature according to the fused feature, the event features, and the event source embedding.
5. The method for generating a clear enhanced image by fusing an image and events under low illumination according to claim 1, wherein The performing cross-scale fusion on the new image features and new event features of different resolutions respectively to obtain new comprehensive image features and new comprehensive event features includes: Input the new image features and new event features with different resolutions into a cross-scale feature fusion module, and the cross-scale feature fusion module outputs comprehensive image features and comprehensive event features, where the cross-scale feature fusion module includes a convolutional layer and a Transformer block; At each resolution scale, extract local information and global information through the convolutional layer and the Transformer block, align features at different resolution scales through a sampling operation, and fuse features at different resolution scales through a convolutional layer.
6. The method for generating a clear enhanced image by fusing an image and events under low illumination according to claim 1, wherein The cross-modal fusion of the comprehensive image features and the comprehensive event features to obtain fused comprehensive features, and generating the image with enhanced clarity of the image data based on the fused comprehensive features includes: Input the comprehensive image features and the comprehensive event features into a cross-modal fusion decoder, and the cross-modal fusion decoder outputs the image with enhanced clarity, where the cross-modal fusion decoder includes a convolutional layer, an activation function, and a Swin-CNN; The convolutional layer and the activation function preprocess the comprehensive image features and the comprehensive event features, and the preprocessed features are input into a first processing channel and a second processing channel; The first processing channel obtains enhanced features through FFT transformation, a Gaussian low-pass filter, an FFC block, and inverse FFT transformation, and the second processing channel is processed through a Swin-CNN block and a convolutional layer, and the features processed by the second processing channel are fused with the enhanced features to obtain fused comprehensive features; The fused comprehensive features are processed through a Swin-CNN and a convolutional layer to obtain the image with enhanced clarity of the image data.
7. The method for generating a clear enhanced image by fusing an image and events under low illumination according to claim 1, characterized in that It also includes: Construct a loss function and use the loss function to guide the training of the model, where the loss function includes a frequency loss, a reconstruction loss, and a total loss. The frequency loss is determined by calculating the norm in the frequency domain, the reconstruction loss is determined by combining the pixel-level error and the structural similarity index, and the total loss is determined by weighting the frequency loss and the reconstruction loss.
8. An apparatus for generating a clear enhanced image by fusing an image and events under low illumination, characterized in that, It includes: An acquisition module for acquiring image data and event data under low illuminance conditions; An extraction module for extracting multi-resolution image features in the image data and multi-resolution event features in the event data. For the multi-resolution image features and multi-resolution event features, new image features and new event features are cross-modally selected at the same resolution scale; A fusion module for respectively performing cross-scale fusion on the new image features and new event features with different resolutions to obtain new comprehensive image features and new comprehensive event features, performing cross-modal fusion on the comprehensive image features and the comprehensive event features to obtain fused comprehensive features, and generating the image with enhanced clarity of the image data based on the fused comprehensive features.
9. An electronic device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the method for fusing images and events under low illuminance to generate an image with enhanced clarity according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instruction is executed, it implements the method for fusing an image and events under low illumination to generate a clear enhanced image according to any one of claims 1-7.
Citation Information
Patent Citations
Multi-modal road scene target detection method based on images and events
CN116453014A
Super-resolution image reconstruction method based on event camera
CN117437125A
Data processing method and device, equipment and medium
CN118262202A
Fuzzy video super-resolution method based on event data driving
CN119850420A
Event camera and Transform-UNet combined video denoising method in low-light environment
CN120047345A
Cited By
End-to-end real-time small target unmanned aerial vehicle detection method based on event camera
CN120564091A