Low-illumination image enhancement method of noise adaptive network based on signal-to-noise ratio guidance

By designing a signal-to-noise ratio-guided noise adaptive network SNA-Net under low light conditions, combining CNN and Transformer models, the problem of difficulty in filtering noise and redundant information in low-light images is solved, and a clear image enhancement effect is achieved.

CN120163720AActive Publication Date: 2025-06-17SOUTHWEST UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510214699.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-17
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

Under low light conditions, the prior art is difficult to effectively filter out noise and redundant information, resulting in poor image enhancement effect.

Method used

A noise adaptive network SNA-Net based on signal-to-noise ratio guidance is designed, combining convolutional neural network CNN and Transformer models to adaptively filter noise and redundant information through the signal-to-noise ratio graph calculation module and feature fusion module.

Benefits of technology

Clear image enhancement is achieved under low light conditions, using information effectively, filtering out irrelevant information, and retaining feature representations of key information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163720A_ABST
    Figure CN120163720A_ABST
Patent Text Reader

Abstract

The invention discloses a low-illumination image enhancement method based on a signal-to-noise ratio guided noise adaptive network. The method comprises the following steps: 1, constructing a low-illumination image enhancement system based on an SNA-Net network; 2, an image acquisition module acquires a low-illumination image; 3, an input layer of the SNA-Net obtains a low-illumination image and transmits the low-illumination image to a CNN encoder, an SNA encoder and a signal-to-noise ratio map calculation module; 4, a signal-to-noise ratio diagram calculation module calculates the low-illumination image to obtain a signal-to-noise ratio diagram; 5, the CNN encoder performs short-distance encoding on the low-illumination image to obtain short-distance features; the SNA encoder performs long-distance encoding on the low-illumination image to obtain a long-distance feature; 6, performing convolution residual operation on the short-distance features by the convolution residual block to obtain convolution residual data; the feature fusion module performs feature fusion operation on the short-distance features and the long-distance features to obtain fused features; 7, the CNN decoder carries out decoding operation on the convolution residual data and the fusion features, and an enhanced image is obtained. The method has the effect that clear image enhancement can be carried out under the low-light condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image enhancement, and particularly to a low-light image enhancement method based on a signal-to-noise ratio-guided noise adaptive network. Background Art

[0002] Images taken under low-light conditions are affected by poor visibility. On the one hand, this affects the quality of people's visual perception; on the other hand, it leads to a decline in the performance of other high-level visual tasks, such as object detection, image segmentation, and recognition.

[0003] There have been many advanced methods for low-light image enhancement. These methods can generally be divided into non-learning-based methods and learning-based methods. In the early stage, more methods only focused on enhancing brightness, contrast, and color factors. In recent years, some works have begun to consider the noise problem in the enhancement process by exploring the signal-to-noise ratio prior to guide the model to focus on different regions of the image. Although noise is still a difficult problem to overcome, considering the information differences in different illumination regions, corresponding methods should be considered for image enhancement. As Figure 1 shown, regions with sufficient illumination (green regions) tend to have more information, and local information is sufficient for local LLIE. In contrast, regions with extremely low illumination (red regions) have less information even though they are dominated by noise, and non-local information is required to effectively enhance the image.

[0004] Due to the insufficient information in the noise regions with extremely low signal-to-noise ratio, and the fact that methods based on convolutional neural network CNN are not good at dealing with long-range dependencies, relying solely on the local learning paradigm based on CNN is not sufficient to reconstruct high-quality images. Therefore, it is beneficial to use a method based on the deep learning model Transformer with self-attention mechanism to capture the non-local information in these regions. At the same time, in regions with relatively high signal-to-noise ratio, local information is already sufficient, and the high computational cost of traditional transformers limits their wide application, and using CNN is sufficient to achieve an ideal reconstruction effect. Therefore, a hybrid model that combines the advantages of CNN and transformers brings new possibilities for low-light enhancement. However, there are still obstacles in designing an effective attention mechanism because the standard transformer uses dense attention calculations, which introduce noise interactions. Recently, to solve this problem, several sparse transformer methods have emerged. Methods based on the Top-k selection attention mechanism and the adaptive sparse attention mechanism have made efforts in this regard. Unfortunately, there is currently no method specifically for the low-light enhancement task that can efficiently filter out redundant information.

[0005] Disadvantages of the prior art: Transformers are of great significance in low-light image enhancement. Low-light images have certain noise, especially in extremely low-light areas. Since Transformers usually calculate the self-attention scores of all available tokens, it is difficult for Transformer-based low-light image enhancement methods to avoid the interference of noise, which is not conducive to clear image enhancement under low-light conditions. Summary of the Invention

[0006] A low-light image enhancement method based on a signal-to-noise ratio-guided noise adaptive network provided by the present invention can perform clear image enhancement under low-light conditions.

[0007] To achieve the above object, a key aspect of a low-light image enhancement method based on a signal-to-noise ratio-guided noise adaptive network provided by the present invention includes the following steps:

[0008] Step 1: Construct a low-light image enhancement system based on a signal-to-noise ratio-guided noise adaptive network. The low-light image enhancement system is provided with an image acquisition module and a signal-to-noise ratio-guided noise adaptive network SNA-Net. The signal-to-noise ratio-guided noise adaptive network SNA-Net is provided with an input layer. The output end of the input layer is connected to the input ends of a CNN encoder, an SNA encoder, and a signal-to-noise ratio map calculation module. The output ends of the CNN encoder, SNA encoder, and signal-to-noise ratio map calculation module are all connected to the input end of a feature fusion module. The output end of the CNN encoder is also connected to the input end of a CNN decoder through a convolutional residual block. The input end of the CNN decoder is also connected to the output end of the feature fusion module;

[0009] Step 2: The image acquisition module acquires a low-light image I and transmits it to the signal-to-noise ratio-guided noise adaptive network SNA-Net;

[0010] Step 3: The input layer of the signal-to-noise ratio-guided noise adaptive network SNA-Net acquires the low-light image I and transmits it to the CNN encoder, SNA encoder, and signal-to-noise ratio map calculation module;

[0011] Step 4: The signal-to-noise ratio map calculation module calculates the low-light image I to obtain a corresponding signal-to-noise ratio map S, and then transmits the signal-to-noise ratio map S to the feature fusion module and the SNA encoder;

[0012] Step 5: The CNN encoder performs short-distance encoding on the low-light image I to capture the local information of the low-light image I, obtaining short-distance feature data a i , and transmits it to the feature fusion module and the convolutional residual block;

[0013] The SNA encoder uses the signal-to-noise ratio map S as prior knowledge to calculate the self-attention score, and then performs long-distance encoding on the low-light image I to filter out noise and redundant information, obtaining long-distance feature data b i , and transmits it to the feature fusion module;

[0014] Step 6: The convolutional residual block performs a convolutional residual operation on the short-distance feature data a i , obtaining convolutional residual data d, and transmits it to the CNN decoder;

[0015] The feature fusion module performs a feature fusion operation on the short-distance feature data a i and the long-distance feature data b i , obtaining fused feature data c i , and transmits it to the CNN decoder;

[0016] Step 7: The CNN decoder performs a decoding operation on the convolutional residual data d and the fused feature data c i , obtaining an enhanced image

[0017] Through the above design, in the SNA-Net network, considering the signal-to-noise ratio changes in different regions, CNN is used for short-distance encoding in the CNN encoder, and Transformer is used for long-distance encoding in the SNA encoder. To better integrate the feature information of both, a signal-to-noise ratio-guided feature fusion module SGFF is constructed to guide the integration process. The feature fusion module SGFF adaptively fuses features of two parallel encoder CNN blocks and SNA blocks through the signal-to-noise ratio map S, and then sends these features to the decoder through residual connections to generate an enhanced image.

[0018] Through the above design, the present invention can not only perform clear image enhancement under low-light conditions, but also use information most effectively, filter out irrelevant information, and retain the feature representation of key information as much as possible.

[0019] Preferably: In the step 1, the CNN encoder is provided with a first CNN encoding block, a second CNN encoding block, a third CNN encoding block, and a fourth CNN encoding block connected in sequence, the SNA encoder is provided with a first SNA block, a second SNA block, a third SNA block, and a fourth SNA block connected in sequence, the feature fusion module is provided with a first feature fusion block, a second feature fusion block, a third feature fusion block, and a fourth feature fusion block, and the CNN decoder is provided with a first CNN decoding block, a second CNN decoding block, a third CNN decoding block, and a fourth CNN decoding block connected in sequence;

[0020] The output ends of the first CNN encoding block and the first SNA block are connected to the input end of the first feature fusion block, and the output end of this first feature fusion block is connected to the input end of the fourth CNN decoding block;

[0021] The output ends of the second CNN encoding block and the second SNA block are connected to the input end of the second feature fusion block, and the output end of this second feature fusion block is connected to the input end of the third CNN decoding block;

[0022] The output ends of the third CNN encoding block and the third SNA block are connected to the input end of the third feature fusion block, and the output end of this third feature fusion block is connected to the input end of the second CNN decoding block;

[0023] The output ends of the fourth CNN encoding block and the fourth SNA block are connected to the input end of the fourth feature fusion block, and the output end of this fourth feature fusion block is connected to the input end of the first CNN decoding block.

[0024] Preferably: in the step 4, first calculate the grayscale image I corresponding to the low-light image I g , and then calculate the signal-to-noise ratio map S according to the following formula, S ∈ R H×W ;

[0025]

[0026] wherein, Denoise() represents the denoising operation using a mean filter, abs represents the absolute value, N ∈ R H×W estimated noise map.

[0027] Preferably: the first CNN encoding block, the second CNN encoding block, the third CNN encoding block, and the fourth CNN encoding block have the same structure and the same working logic, and are all provided with a 3×3 convolutional layer and a downsampling layer, and a LeakyReLU activation function is built in the output end of this 3×3 convolutional layer.

[0028] Each stage in the CNN encoder is implemented by a CNN block and a single convolutional layer for downsampling. The CNN block consists of a 3×3 convolutional layer and a LeakyReLU activation function, and is used to capture local information.

[0029] Preferably: the first SNA block, the second SNA block, the third SNA block, and the fourth SNA block have the same structure and the same working logic, and are all provided with a noise adaptive self-attention module NASA and a dual-domain refinement feedforward network DFRN;

[0030] In each SNA block, the Noise Adaptive Self-Attention module NASA calculates the self-attention score Attn of the corresponding SNA block by using the signal-to-noise ratio map S as prior knowledge. Then, the long-distance encoding of the low-light image I or the output data of the previous SNA block is performed through the self-attention score Attn to obtain the self-attention feature of the current SNA block. Subsequently, the Dual-Domain Refinement Feed-Forward Network DRFN further refines the self-attention feature to obtain the long-distance feature data b of the current SNA block. i 。

[0031] The Noise Adaptive Self-Attention module NASA can eliminate the possibility of loss of information integrity, retain valid information as much as possible, and filter out irrelevant features to the greatest extent. The Dual-Domain Refinement Feed-Forward Network DRFN can refine the features in the spatial and frequency domains to eliminate potential redundant information, that is, enhance the most useful features for image restoration in both domains while suppressing potential redundant features.

[0032] The Dual-Domain Refinement Feed-Forward Network DRFN suppresses redundant information in the channel dimension, while the Noise Adaptive Self-Attention module NASA suppresses the noise interaction in the irrelevant regions in the spatial dimension. The complementarity of these two components enables SNA-Net to suppress irrelevant features to a certain extent while obtaining the most informative feature representation.

[0033] Preferably, the Noise Adaptive Self-Attention module NASA obtains the self-attention feature of the current SNA block through the following steps:

[0034] Step A1: The first normalization layer in the Noise Adaptive Self-Attention module NASA obtains the low-light image I or the output data of the previous SNA block, generates tensor data, and then passes it to the query matrix Q encoding unit, the key matrix K encoding unit, the value matrix V encoding unit, and the first addition unit.

[0035] Step A2: The query matrix encoding unit encodes the tensor data through the first 1×1 convolutional layer and the first 3×3 depth convolutional layer to generate the query matrix Q, and passes it to the first multiplication unit.

[0036] The key matrix K encoding unit encodes the tensor data through the second 1×1 convolutional layer and the second 3×3 depth convolutional layer to generate the key matrix K, and after transposing the key matrix K, passes it to the first multiplication unit.

[0037] The value matrix V encoding unit encodes the tensor data through the third 1×1 convolutional layer and the third 3×3 depth convolutional layer to generate the value matrix V, and passes it to the second multiplication unit.

[0038] The expressions for generating the query matrix Q, the key matrix K, and the value matrix V are as follows:

[0039]

[0040] where, W d represents a 1×1 pointwise convolution, W p represents a 3×3 depthwise convolution, and X represents tensor data;

[0041] Step A3: The first multiplication unit multiplies the query matrix Q and the transposed key matrix K T to obtain a first multiplication matrix, and passes it to the sparse self-attention branch and the dense self-attention branch;

[0042] Step A4: The dense self-attention branch calculates first attention data through a first Softmax function based on the first multiplication matrix, and passes it to the second addition unit;

[0043] The sparse self-attention branch uses the normalized signal-to-noise ratio map S′ as a mask to filter out the regions in the first multiplication matrix where the signal-to-noise ratio values are lower than the signal-to-noise ratio threshold, and then calculates the scores of the regions where the signal-to-noise ratio values are higher than or equal to the signal-to-noise ratio threshold through a second Softmax function to obtain second attention data, and passes it to the second addition unit;

[0044] Step A5: The second addition unit weights the first attention data with a first attention weight w1, weights the second attention data with a second attention weight w2, adds the weighted first attention score and the second attention score, and then multiplies the sum by the value matrix V through a second multiplication unit to obtain the self-attention score Attn, and the expression is as follows:

[0045] Attn = w1 * D attn + w2 * S attn

[0046] D attn = Softmax(QK T / α)V

[0047] S attn = Softmax(QK T / α+(1 - S′)σ)V

[0048] where, α is a learnable scaling parameter used to control the dot product size of the query matrix Q and the key matrix K; σ is a small negative scalar -1e9, D attn is the dense self-attention score, S attnis the sparse self-attention score; w1 and w2 are two normalized weights used to adaptively combine the two branches, and * is the multiplication operation. A positional encoding with learnable parameters is added at the end of the noise adaptive self-attention module to produce the final output. This design can ensure that while filtering out irrelevant feature interactions such as noise, sufficient information features are retained. That is, the model can well control the sparsity degree of the input tokens of a specific task.

[0049] Step A6: The second multiplication unit passes the self-attention score Attn to the fourth 1×1 convolutional layer for convolution operation, and then the convolution result is element-wise added to the tensor data through the first addition unit to obtain the self-attention feature of the current SNA block, and is passed to the dual-domain refinement feed-forward network DRFN.

[0050] To weaken the negative impact of noise regions or other irrelevant information and prevent the possibility of excessive sparsity, the noise adaptive self-attention module NASA is proposed. The noise adaptive self-attention module NASA uses a channel-level self-attention mechanism to calculate the self-attention score in the channel dimension, thereby reducing complexity. NASA consists of a dense branch and a sparse branch, and these two branches can be adaptively fused. The sparse branch filters out the tokens from the low signal-to-noise ratio regions to avoid noise interference, while the dense branch ensures the integrity of key information.

[0051] The sparse self-attention branch uses the normalized signal-to-noise ratio map as a mask to filter out the low signal-to-noise ratio regions and only calculates the scores for the relatively high signal-to-noise ratio regions. This sparse method can filter out redundant information and irrelevant features. At the same time, considering the possibility of excessive sparsity, another dense self-attention branch is introduced. This branch uses the standard softmax attention to retain complete information. NASA adaptively obtains features from the two branches and propagates the information flow through the network.

[0052] Preferably: The dual-domain refinement feed-forward network DRFN further refines the self-attention feature through the following steps:

[0053] Step B1: The second normalization layer in the dual-domain refinement feed-forward network DRFN normalizes the self-attention feature, and then the normalized feature is convolved by the fifth 1×1 convolutional layer to obtain the fifth convolution data, and is passed to the sixth 1×1 convolutional layer, the fourth 3×3 depth convolutional layer, and the fast Fourier transform unit FFT;

[0054] The expression for obtaining the fifth convolution data is as follows:

[0055] Y′ = W d (LN(Y))

[0056] Among them, Y is the self-attention feature, LN is layer normalization, Y' is the fifth convolutional data, and W d represents a 1×1 pointwise convolution;

[0057] Step B2: The sixth 1×1 convolutional layer performs a convolution operation on the fifth convolutional data, and then performs a depth convolution operation through the fifth 3×3 depth convolutional layer to obtain the fifth depth convolutional data, which is then passed to the first element-wise multiplication unit;

[0058] The expression for obtaining the fifth depth convolutional data is as follows:

[0059] Y″ = W p W d (Y')[[]]

[0060] where W p represents a 3×3 depth convolution, and Y″ represents the fifth depth convolutional data;

[0061] The fourth 3×3 depth convolutional layer performs a depth convolution operation on the fifth convolutional data, and then performs activation through the GELU activation function to obtain the first activation data, which is then passed to the element-wise multiplication unit;

[0062] The expression for obtaining the first activation data is as follows:

[0063]

[0064] where is the GELU activation function, and Y″′ represents the first activation data;

[0065] Step B3: The fast Fourier transform unit FFT performs a fast Fourier transform on the fifth convolutional data to obtain FFT data, and then the FFT data is element-wise multiplied with the weight w of the dual-domain refinement feed-forward network DRFN through the second element-wise multiplication unit, and then an inverse fast Fourier transform is performed through the inverse fast Fourier transform unit IFFT, and then activation is performed through the GEGLU activation function to obtain the second activation data, which is then passed to the connection unit;

[0066] The expression for obtaining the second activation data is as follows:

[0067]

[0068] where is the fast Fourier transform, is the inverse fast Fourier transform, w is the weight of the dual-domain refinement feed-forward network DRFN, Y f is the second activation data, and ζ() represents the GEGLU function;

[0069] The first element-wise multiplication unit multiplies the fifth depth convolution data and the first activation data element-wise to obtain first multiplication data and passes it to the connection unit;

[0070] The expression for obtaining the first multiplication data is as follows:

[0071] Y s = Y″ ⊙ Y″′

[0072] where ⊙ represents element-wise multiplication, and Y s is the first multiplication data;

[0073] Step B4: The connection unit performs a connection operation on the second activation data and the first multiplication data, and then performs a convolution operation through the seventh 1×1 convolution layer to obtain seventh convolution data and passes it to the third addition unit;

[0074] Step B5: The third addition unit adds the self-attention feature and the seventh convolution data element-wise to obtain the long-range feature data b i ;

[0075] The expression for obtaining the long-range feature data b i is as follows:

[0076]

[0077] where represents the connection operation.

[0078] The dual-domain refinement feed-forward network DRFN refines features in the spatial domain and the frequency domain to enhance feature representation, thereby achieving better potential image enhancement. Specifically, in the frequency domain, DRFN uses a learnable global filter to determine which low-frequency and high-frequency information should be retained to restore the underlying clear image. In the spatial domain, a gating mechanism is used to reduce redundant features.

[0079] Preferably: The first feature fusion block, the second feature fusion block, the third feature fusion block, and the fourth feature fusion block have the same structure and the same working logic, and are all provided with a third multiplication unit, a fourth multiplication unit, a fourth addition unit, a global average pooling layer, a first fully-connected layer, a second fully-connected layer, a fifth multiplication unit, a sixth multiplication unit, and a fifth addition unit. The output end of the second fully-connected layer is built-in with a Sigmoid function;

[0080] In step 6, the feature fusion block obtains the corresponding short-range feature data a i and the long-range feature data b i and performs a feature fusion operation to obtain the fusion feature data c i, including the following steps:

[0081] Step C1: The third multiplication unit in the current feature fusion block uses the normalized signal-to-noise ratio map as a mask, multiplies it with the short-distance feature data a i to obtain the third multiplication data, and passes it to the fourth addition unit and the fifth multiplication unit;

[0082] The fourth multiplication unit in the current feature fusion block multiplies the normalized signal-to-noise ratio map with the long-distance feature data b i to obtain the fourth multiplication data, and passes it to the fourth addition unit and the sixth multiplication unit;

[0083] Step C2: The fourth addition unit element-wise adds the third multiplication data and the fourth multiplication data, then performs global average pooling operation through the global average pooling layer, and then performs fully connected operations through the first fully connected layer and the second fully connected layer. After normalization by the Sigmoid function, the short-distance feature fusion weight w a and the long-distance feature fusion weight w d are obtained, and the short-distance feature fusion weight w a is passed to the fifth multiplication unit, and the long-distance feature fusion weight w d is passed to the sixth multiplication unit;

[0084] Step C3: The fifth multiplication unit weights the third multiplication data by the short-distance feature fusion weight w a to obtain the fifth multiplication data, and passes it to the fifth addition unit;

[0085] The sixth multiplication unit weights the fourth multiplication data by the long-distance feature fusion weight w d to obtain the sixth multiplication data, and passes it to the fifth addition unit;

[0086] Step C4: The fifth addition unit element-wise adds the fifth multiplication data and the sixth multiplication data to obtain the feature fusion data c i of the current feature fusion block, and the expression is as follows:

[0087] c i = SGFF(a i × (1 - S′) + b i × S′)

[0088] where SGFF represents the feature fusion operation, S′ represents the normalized signal-to-noise ratio map, and 1 - S′ represents using the normalized signal-to-noise ratio map as a mask.

[0089] Simply concatenating the features of two parallel encoders may lead to an over - contribution of unimportant information. To address this issue and naturally combine the features of the two encoders, a signal - to - noise ratio guided feature fusion module SGFF is developed to dynamically integrate the information flow. Since the signal - to - noise ratio map reflects different noise levels in different image regions, SGFF can adaptively combine the long - range features from the SNA block and the short - range features from the CNN block.

[0090] Specifically, the feature fusion module uses the normalized signal - to - noise ratio map as an interpolation weight to selectively fuse features, leveraging their respective capabilities in global and local feature extraction to generate two merged features. Subsequently, the features are further processed by applying global average pooling. Next, through two fully - connected layers, the dimension is first reduced and then increased to enhance the expressiveness of the information. Finally, after softmax normalization, two weights are obtained, which are used to weight and sum the features, thus obtaining the final fused features.

[0091] The beneficial effects of the present invention: A novel hybrid CNN - Transformer dual - encoder model, named signal - to - noise ratio guided noise - adaptive network SNA - Net, is proposed. It can adapt to the signal - to - noise ratio changes in different illumination regions and utilize the respective advantages of the convolutional neural network CNN and the deep - learning model Transformer based on the self - attention mechanism to meet the enhancement requirements of different regions of the image. SNA - Net mainly includes two components: the noise - adaptive self - attention module NASA and the dual - domain refinement feed - forward network DRFN. NASA adaptively calculates the attention scores from the dense and sparse branches. The sparse branch filters out the negative marker interactions from low signal - to - noise ratio regions guided by the signal - to - noise ratio map, while the dense branch ensures the integrity of the information. At the same time, DRFN eliminates feature redundancy in the dual domain to improve the recovery of the potential clear image. In addition, to better integrate the feature information from CNN and Transformer, a signal - to - noise ratio guided feature fusion module SGFF is designed. Brief Description of the Drawings

[0092] Figure 1 It is a schematic diagram of the structure of the signal - to - noise ratio guided noise - adaptive network SNA - Net in the embodiment;

[0093] Figure 2 It is a schematic diagram of the structure of the noise - adaptive self - attention module NASA in the embodiment;

[0094] Figure 3 It is a schematic diagram of the structure of the dual - domain refinement feed - forward network DRFN in the embodiment;

[0095] Figure 4 It is a schematic diagram of the structure of the feature fusion block in the embodiment;

[0096] Figure 5 Comparison charts of the results of different models in the embodiments on LOLv2-Real (upper) and LOLv2-Synthetic (lower);

[0097] Figure 6 Comparison chart of the quantitative results of different models in the embodiments on the reference-free dataset. Detailed implementation manners

[0098] The present invention will be further described in detail below with reference to the accompanying drawings and specific examples. The following embodiments or drawings are used to illustrate the present invention, but not to limit the scope of the present invention.

[0099] As shown in FIG. 1: A low-light image enhancement method based on a signal-to-noise ratio-guided noise adaptive network includes the following steps:

[0100] Step 1: Construct a low-light image enhancement system based on a signal-to-noise ratio-guided noise adaptive network. The low-light image enhancement system is provided with an image acquisition module and a signal-to-noise ratio-guided noise adaptive network SNA-Net. The signal-to-noise ratio-guided noise adaptive network SNA-Net is provided with an input layer. The output end of the input layer is connected to the input ends of a CNN encoder, an SNA encoder, and a signal-to-noise ratio map calculation module. The output ends of the CNN encoder, the SNA encoder, and the signal-to-noise ratio map calculation module are all connected to the input end of a feature fusion module. The output end of the CNN encoder is also connected to the input end of a CNN decoder through a convolutional residual block. The input end of the CNN decoder is also connected to the output end of the feature fusion module;

[0101] Step 2: The image acquisition module acquires a low-light image I and transmits it to the signal-to-noise ratio-guided noise adaptive network SNA-Net;

[0102] Step 3: The input layer of the signal-to-noise ratio-guided noise adaptive network SNA-Net acquires the low-light image I and transmits it to the CNN encoder, the SNA encoder, and the signal-to-noise ratio map calculation module;

[0103] Step 4: The signal-to-noise ratio map calculation module calculates the low-light image I to obtain a corresponding signal-to-noise ratio map S, and then transmits the signal-to-noise ratio map S to the feature fusion module and the SNA encoder;

[0104] Step 5: The CNN encoder performs short-distance encoding on the low-light image I to capture local information of the low-light image I, and obtains short-distance feature data a i , and transmits it to the feature fusion module and the convolutional residual block;

[0105] The SNA encoder uses the signal-to-noise ratio map S as prior knowledge to calculate the self-attention scores, and then performs long-range encoding on the low-light image I to filter out noise and redundant information, obtaining long-range feature data b i , and transmits it to the feature fusion module;

[0106] Step 6: The convolutional residual block performs a convolutional residual operation on the short-range feature data a i , obtaining convolutional residual data d, and transmits it to the CNN decoder;

[0107] The feature fusion module performs a feature fusion operation on the short-range feature data a i and the long-range feature data b i , obtaining fused feature data c i , and transmits it to the CNN decoder;

[0108] Step 7: The CNN decoder performs a decoding operation on the convolutional residual data d and the fused feature data c i , obtaining the enhanced image

[0109] For the low-light image I obtained by the image acquisition module, first, a signal-to-noise ratio map S is obtained using a method without learning. Then, I is input into two parallel encoders composed of a CNN block and an SNA block. Subsequently, the signal-to-noise ratio map is used to guide the fusion process between the CNN block and the SNA block, and the fused features are transmitted to the corresponding layers of the decoder through skip connections to generate the enhanced image

[0110] This embodiment adopts the Charbonnier loss and the perceptual loss to train the SNA-Net network, and the overall loss function is:

[0111]

[0112] where, I′ is the real data, ε is set to 10 -3 , λ is a hyperparameter, ‖‖2 represents the L2 norm, and ‖‖1 represents the L1 norm.

[0113] In step 4, first, the grayscale image I corresponding to the low-light image I is calculated g , and then the signal-to-noise ratio map S is calculated according to the following formula, S ∈ R H×W ;

[0114]

[0115] Among them, Denoise() represents the denoising operation using a mean filter, abs represents the absolute value, and N ∈ R H×W The estimated noise map.

[0116] The CNN encoder is provided with a first CNN encoding block, a second CNN encoding block, a third CNN encoding block, and a fourth CNN encoding block connected in sequence. The SNA encoder is provided with a first SNA block, a second SNA block, a third SNA block, and a fourth SNA block connected in sequence. The feature fusion module is provided with a first feature fusion block, a second feature fusion block, a third feature fusion block, and a fourth feature fusion block. The CNN decoder is provided with a first CNN decoding block, a second CNN decoding block, a third CNN decoding block, and a fourth CNN decoding block connected in sequence;

[0117] The output ends of the first CNN encoding block and the first SNA block are connected to the input end of the first feature fusion block, and the output end of this first feature fusion block is connected to the input end of the fourth CNN decoding block;

[0118] The output ends of the second CNN encoding block and the second SNA block are connected to the input end of the second feature fusion block, and the output end of this second feature fusion block is connected to the input end of the third CNN decoding block;

[0119] The output ends of the third CNN encoding block and the third SNA block are connected to the input end of the third feature fusion block, and the output end of this third feature fusion block is connected to the input end of the second CNN decoding block;

[0120] The output ends of the fourth CNN encoding block and the fourth SNA block are connected to the input end of the fourth feature fusion block, and the output end of this fourth feature fusion block is connected to the input end of the first CNN decoding block.

[0121] The first CNN encoding block, the second CNN encoding block, the third CNN encoding block, and the fourth CNN encoding block have the same structure and the same working logic, and are all provided with a 3×3 convolutional layer and a downsampling layer. The output end of this 3×3 convolutional layer is built-in with a LeakyReLU activation function.

[0122] The first SNA block, the second SNA block, the third SNA block, and the fourth SNA block have the same structure and the same working logic, and are all provided with a Noise Adaptive Self-Attention module NASA and a Dual-Domain Refinement Feed-Forward Network DRFN;

[0123] In each SNA block, the noise adaptive self-attention module NASA calculates the self-attention score Attn of the corresponding SNA block by using the signal-to-noise ratio map S as prior knowledge. Then, the long-distance encoding of the low-light image I or the output data of the previous SNA block is performed through the self-attention score Attn to obtain the self-attention feature of the current SNA block. Subsequently, the self-attention feature is further refined by the dual-domain refinement feed-forward network DRFN to obtain the long-distance feature data b of the current SNA block i .

[0124] As shown in Figure 2: The noise adaptive self-attention module NASA obtains the self-attention feature of the current SNA block through the following steps:

[0125] Step A1: The first normalization layer in the noise adaptive self-attention module NASA obtains the output data of the low-light image I or the previous SNA block, generates tensor data, and then passes it to the query matrix encoding Q unit, the key matrix K encoding unit, the value matrix V encoding unit, and the first addition unit;

[0126] Step A2: The query matrix encoding unit encodes the tensor data through the first 1×1 convolutional layer and the first 3×3 depth convolutional layer to generate the query matrix Q, and passes it to the first multiplication unit;

[0127] The key matrix K encoding unit encodes the tensor data through the second 1×1 convolutional layer and the second 3×3 depth convolutional layer to generate the key matrix K. After transposing the key matrix K, it is passed to the first multiplication unit;

[0128] The value matrix V encoding unit encodes the tensor data through the third 1×1 convolutional layer and the third 3×3 depth convolutional layer to generate the value matrix V, and passes it to the second multiplication unit;

[0129] The expressions for generating the query matrix Q, the key matrix K, and the value matrix V are as follows:

[0130]

[0131] where W d represents a 1×1 pointwise convolution, W p represents a 3×3 depth convolution, and X represents tensor data;

[0132] Step A3: The first multiplication unit multiplies the query matrix Q and the transposed key matrix K T to obtain the first multiplication matrix, and passes it to the sparse self-attention branch and the dense self-attention branch;

[0133] Step A4: The dense self-attention branch calculates first attention data through the first Softmax function according to the first multiplication matrix, and transmits it to the second addition unit;

[0134] The sparse self-attention branch uses the normalized signal-to-noise ratio map S′ as a mask to filter out the regions in the first multiplication matrix where the signal-to-noise ratio is lower than the signal-to-noise ratio threshold, and then calculates the scores of the regions where the signal-to-noise ratio is higher than or equal to the signal-to-noise ratio threshold through the second Softmax function to obtain second attention data, and transmits it to the second addition unit;

[0135] Step A5: The second addition unit weights the first attention data with the first attention weight w1, weights the second attention data with the second attention weight w2, adds the weighted first attention score and the second attention score, and then multiplies the result by the value matrix V through the second multiplication unit to obtain the self-attention score Attn. The expression is as follows:

[0136] Attn = w1 * D attn + w2 * S attn

[0137] D attn = Softmax(QK T / α)V

[0138] S attn = Softmax(QK T / α + (1 - S′)σ)V

[0139] where α is a learnable scaling parameter used to control the dot product size of the query matrix Q and the key matrix K; σ is a small negative scalar -1e9, D attn is the dense self-attention score, S attn is the sparse self-attention score; w1 and w2 are two normalized weights used to adaptively combine the two branches, and * is the multiplication operation.

[0140] The noise adaptive self-attention module NASA divides the channels into multiple heads and learns separate attention maps in parallel. Therefore, it can be ensured that the long-distance attention comes from the image regions with relatively high signal-to-noise ratio.

[0141] Step A6: The second multiplication unit transmits the self-attention score Attn to the fourth 1×1 convolutional layer for convolution operation, and then adds the convolution result to the tensor data element-wise through the first addition unit to obtain the self-attention feature of the current SNA block, and transmits it to the dual-domain refinement feed-forward network DRFN.

[0142] For the given input feature map and the signal-to-noise ratio map First, the size of the signal-to-noise ratio map S is adjusted pixel by pixel to match the input feature map I, and then it is normalized. In the self-attention score calculation, each pixel acts as a mask, which can effectively suppress the influence of image regions with very low signal-to-noise ratio during the enhancement process. The value of each pixel is:

[0143]

[0144] where T represents the threshold of the signal-to-noise ratio map, S * represents the signal-to-noise ratio map after pixel-by-pixel adjustment, and S′ represents the normalized signal-to-noise ratio map.

[0145] First, layer normalization is used to generate a tensor, and then the cross-channel local spatial context is encoded through cascaded 1×1 convolution and 3×3 depth convolution to generate the query matrix Q, the key matrix K, and the value matrix V. Then, Q and K are reshaped so that the size of the transposed attention map obtained through their dot product interaction is R C×C , rather than the size R of the regular huge attention map HW×HW . Since not all query matrices Q are closely related to the key matrix K, it is not the best choice to calculate attention scores for all tokens during the enhancement process. Using signal-to-noise ratio map-guided sparse self-attention score calculation may be a reasonable solution, which filters out the irrelevant information in the regions with extremely low signal-to-noise ratio and propagates the most useful information flow.

[0146] As Figure 3 shown: The dual-domain refinement feed-forward network DRFN further refines the self-attention features through the following steps:

[0147] Step B1: The second normalization layer in the dual-domain refinement feed-forward network DRFN normalizes the self-attention features, and then the normalized features are convolved by the fifth 1×1 convolutional layer to obtain the fifth convolutional data, which is passed to the sixth 1×1 convolutional layer, the fourth 3×3 depth convolutional layer, and the fast Fourier transform unit FFT;

[0148] The expression for obtaining the fifth convolutional data is as follows:

[0149] Y′ = W d (LN(Y))

[0150] where Y is the self-attention feature, LN is layer normalization, Y′ is the fifth convolutional data, and W d represents a 1×1 pointwise convolution;

[0151] Step B2: The sixth 1×1 convolutional layer performs a convolution operation on the fifth convolutional data, and then performs a depth convolution operation through the fifth 3×3 depth convolutional layer to obtain fifth depth convolutional data, which is passed to the first element-wise multiplication unit;

[0152] The expression for obtaining the fifth depth convolutional data is as follows:

[0153] Y″ = W p W d (Y′)

[0154] where W p represents a 3×3 depth convolution, and Y″ represents the fifth depth convolutional data;

[0155] The fourth 3×3 depth convolutional layer performs a depth convolution operation on the fifth convolutional data, and then performs activation through the GELU activation function to obtain first activation data, which is passed to the element-wise multiplication unit;

[0156] The expression for obtaining the first activation data is as follows:

[0157]

[0158] where is the GELU activation function, and Y″′ represents the first activation data;

[0159] Step B3: The Fast Fourier Transform unit FFT performs a fast Fourier transform on the fifth convolutional data to obtain FFT data, and then the FFT data is element-wise multiplied with the weight w of the Dual-Domain Refinement Feed-Forward Network DRFN through the second element-wise multiplication unit, and then an inverse fast Fourier transform is performed through the Inverse Fast Fourier Transform unit IFFT, and then activation is performed through the GEGLU activation function to obtain second activation data, which is passed to the connection unit;

[0160] The expression for obtaining the second activation data is as follows:

[0161]

[0162] where is the fast Fourier transform, is the inverse fast Fourier transform, w is the weight of the Dual-Domain Refinement Feed-Forward Network DRFN, Y f is the second activation data, and ζ() represents the GEGLU function;

[0163] The first element-wise multiplication unit element-wise multiplies the fifth depth convolutional data and the first activation data to obtain first multiplication data, which is passed to the connection unit;

[0164] The expression for obtaining the first multiplication data is as follows:

[0165] Y s = Y″⊙Y″′

[0166] where ⊙ represents element-wise multiplication, and Y s is the first multiplication data;

[0167] Step B4: The connection unit performs a connection operation on the second activation data and the first multiplication data, and then performs a convolution operation through the seventh 1×1 convolution layer to obtain seventh convolution data, and transmits it to the third addition unit;

[0168] Step B5: The third addition unit performs element-wise addition of the self-attention feature and the seventh convolution data to obtain the long-range feature data b of the current SNA block i ;

[0169] Obtain the long-range feature data b of the current SNA block i The expression is as follows:

[0170]

[0171] where represents the connection operation.

[0172] The common fast Fourier transform processes the information of each pixel position separately and plays a very important role in improving feature representation. Therefore, in the image restoration task, designing an effective feed-forward network is crucial for generating features that contribute to reconstructing high-quality images. To further refine the characteristics generated by NASA, a dual-domain refinement feed-forward network DRFN is designed to enhance the flow of effective information in the features. Specifically, the present invention constructs DRFN by performing two parallel operations in the spatial and frequency domains. In the spatial branch, the information flow is adjusted through a gating mechanism, which suppresses features containing redundant information and only allows useful information to propagate further. In the frequency branch, since not all low-frequency and high-frequency information is beneficial to image restoration, a learnable quantization matrix is used as a global filter during the frequency modulation process to adaptively determine which frequency information should be retained.

[0173] As Figure 4 shown: The first feature fusion block, the second feature fusion block, the third feature fusion block, and the fourth feature fusion block have the same structure and the same working logic, and are all provided with a third multiplication unit, a fourth multiplication unit, a fourth addition unit, a global average pooling layer, a first fully connected layer, a second fully connected layer, a fifth multiplication unit, a sixth multiplication unit, and a fifth addition unit. A Sigmoid function is built in at the output end of the second fully connected layer;

[0174] In the step 6, the feature fusion block obtains the corresponding short-range feature data ai and long - distance feature data b i and perform a feature fusion operation to obtain the fused feature data c of the current feature fusion block i , including the following steps:

[0175] Step C1: The third multiplication unit in the current feature fusion block uses the normalized signal - to - noise ratio map as a mask and multiplies it with the short - distance feature data a i to obtain the third multiplication data, and passes it to the fourth addition unit and the fifth multiplication unit;

[0176] The fourth multiplication unit in the current feature fusion block multiplies the normalized signal - to - noise ratio map with the long - distance feature data b i to obtain the fourth multiplication data, and passes it to the fourth addition unit and the sixth multiplication unit;

[0177] Step C2: The fourth addition unit performs element - by - element addition on the third multiplication data and the fourth multiplication data, then performs global average pooling operation through the global average pooling layer, then performs fully - connected operations through the first fully - connected layer and the second fully - connected layer, and then after normalization by the Sigmoid function, obtains the short - distance feature fusion weight w a and the long - distance feature fusion weight w d , and passes the short - distance feature fusion weight w a to the fifth multiplication unit, and passes the long - distance feature fusion weight w d to the sixth multiplication unit;

[0178] Step C3: The fifth multiplication unit weights the third multiplication data by the short - distance feature fusion weight w a to obtain the fifth multiplication data, and passes it to the fifth addition unit;

[0179] The sixth multiplication unit weights the fourth multiplication data by the long - distance feature fusion weight w d to obtain the sixth multiplication data, and passes it to the fifth addition unit;

[0180] Step C4: The fifth addition unit performs element - by - element addition on the fifth multiplication data and the sixth multiplication data to obtain the fused feature data c of the current feature fusion block i , and the expression is as follows:

[0181] c i = SGFF(a i ×(1 - S′)+b i ×S′)

[0182] where SGFF represents the feature fusion operation, S′ represents the normalized signal - to - noise ratio map, and 1 - S′ represents using the normalized signal - to - noise ratio map as a mask.

[0183] Next, specific experiments are conducted to evaluate the performance of the SNR-guided noise adaptive network SNA-Net proposed in the present invention.

[0184] Dataset: In this embodiment, validity tests are carried out on two reference datasets, LOLv2-Real and LOLv2-Synthetic. Among them, the LOLv2-Real dataset contains 689 pairs of low-light / normal-light images for training and 100 pairs for testing. The LOLv2-Synthetic dataset contains 900 pairs of images for training and 100 pairs for testing.

[0185] In addition, this embodiment also tests our method on four no-reference datasets, namely LIME, DICM, MEF, and NPE.

[0186] Implementation details: In this embodiment, the framework of the present invention is implemented through the open-source deep learning framework PyTorch for machine learning and deep learning, and training and testing are carried out on an NVIDIA GeForce RTX 3090 GPU. For SNA-Net, the numbers of the first SNA block, the second SNA block, the third SNA block, and the fourth SNA block are 4, 6, 6, and 8 respectively, and the number of self-attention heads in the SNA blocks at the same level is set to {1, 2, 4, 8}. During training, the Adam optimizer is used to optimize the SNA-Net network model, the batch size is set to 8, the patch size is set to 128, and a total of 600k iterations are performed. For data augmentation, vertical and horizontal flips are randomly applied. Peak signal-to-noise ratio PSNR and structural similarity SSIM are used as evaluation metrics on the reference datasets, and natural image quality evaluator NIQE and blind / no-reference image spatial quality evaluator BRISQUE are used on the no-reference datasets.

[0187] Qualitative results: The present invention is compared with some state-of-the-art technologies SOTA to demonstrate its superiority. As Figure 6 shown, first, color distortion and image blurring occur during the enhancement process of Restormer on the LOLv2-Real dataset. Second, SNR-Net shows serious artificial traces on the no-reference datasets. Third, as Figure 5 shown in the results of LOLv2-Synthetic in Figure 6As shown. Finally, the results of LLformer also show its noise problem. Unfortunately, previous methods have some problems in color distortion, overexposure / underexposure, or image blurring. In contrast, the SNA-Net network can stably enhance low visibility while preserving as much detail information as possible, without causing overexposure / underexposure, and without introducing additional noise or artificial artifacts.

[0188] Quantitative results:

[0189] Signal-to-Noise Ratio Guided Noise Adaptive Network

[0190]

[0191] Table 1 Comparison of Quantitative Results on LOLv2-Real and LOLv2-Synthetic Datasets

[0192] Signal-to-Noise Ratio Guided Noise Adaptive Network

[0193]

[0194] Table 2 Quantitative Results of Different Models on the No-Reference Dataset

[0195] In this part, the quantitative results of all models are presented in Table 1 and Table 2. The performance of the SNA-Net network has been significantly improved on most datasets. Specifically, on the LOLv2-Real dataset, for PSNR, SNA-Net has improved by 0.91dB and 0.05dB compared to SNR-Net and FourLLIE respectively, indicating better improvement in noise suppression. For the structural similarity SSIM, the present invention has also achieved significant improvements on both datasets, indicating that the images enhanced by the SNA-Net network have better perceptual quality. In summary, compared with other methods, all four evaluation metrics used have been improved to a certain extent, proving that the present invention has achieved satisfactory restoration effects in terms of structural details and noise. The quantitative metrics on the no-reference, unpaired dataset also indicate that the SNA-Net network has certain advantages.

[0196] Low-light object detection:

[0197] Low-light image enhancement, as a preprocessing step for advanced vision tasks, plays a crucial role in improving their performance. To this end, the present invention uses the Exdark dataset for low-light object detection experiments. And the effects of using different low-light enhancement algorithms as preprocessing steps on low-light detection performance are compared. The Exdark dataset contains 7,363 images, covering 10 different low-light environments. In this experiment, YOLOv5s is used as the object detection model and combined with various enhancers, which are used as preprocessing modules with fixed parameters. By training the model, the effects of these preprocessing steps on detection performance are evaluated. As Figure 3 shown, the model of the present invention achieves the best performance. Compared with LLformer, it improves by 3.76% in mAR 0.5 and is 4.64% higher than SNR-Net in mAR 0.5:0.95 . These improvements may stem from the adaptive noise filtering of NASA and the refinement of DRFN, which enable the enhancer to recover enhanced images with reduced noise interference and richer details, which is crucial for the detector.

[0198] Ablation experiments:

[0199] To verify the effectiveness of each component in the SNA-Net network, three main ablation studies are considered, which are carried out on the LOLv2-Real dataset by removing three different components from the SNA-Net network framework. In addition, some ablation studies are also considered for two components of the SNA block.

[0200] · Without the long branch: The Transformer encoder branch is removed, so the framework only has a CNN-based encoder.

[0201] · Without the short branch: The CNN encoder branch is removed, so the framework only has a Transformer-based encoder..

[0202] · Without SGFF: The signal-to-noise ratio-guided feature fusion module is removed, and only simple concatenation is used.

[0203] Table 3 shows the performance comparison of low-light object detection on the Exdark dataset preprocessed by different enhancers.

[0204]

[0205] Table 3 Performance of low-light object detection on the Exdark dataset preprocessed by different enhancers

[0206] As shown in Table 4, the complete setup achieved the highest PSNR and superior SSIM values. "w / o S" and "w / o L" demonstrated the effectiveness of the two parallel encoders. "w / 0SGFF" demonstrated that the SGFF module could more finely integrate the two features from the SNA block and the CNN encoding block.

[0207]

[0208] Table 4 Ablation experiment results of SNA-Net

[0209] Effectiveness of NASA: To study the effectiveness of the NASA component, several existing self-attention mechanisms were replaced for comparison. As shown in Table 5, compared with MDTA in Restormer, the PSNR and SSIM increased by 0.71 dB and 0.01 respectively. Comparisons were also made with TKSA and ASSA, and certain performance improvements were obtained. In addition, since the sparse method did not introduce additional parameters, it did not lead to an increase in complexity while improving performance.

[0210]

[0211] Table 5 Ablation experiments for different self-attention mechanisms

[0212] Effectiveness of DRFN: Since not all features contribute positively to restoring clear images, simply applying feature transformation may result in too much redundant information. DREN estimated the information useful for image restoration in the frequency and spatial domains. To prove its effectiveness in the dual domain, ablation studies were conducted by removing the frequency domain and the spatial domain respectively. Specifically, the experimental settings included (1) without DRFN, (2) without the frequency branch, and (3) without the spatial branch. The quantitative comparison results are shown in Table 6. The results indicate that DRFN can select more effective information in the two domains, reduce redundant features, and thus better complement the NASA design.

[0213]

[0214] Table 6 Ablation experiments of the dual-domain refinement feed-forward network

[0215] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A low-light image enhancement method based on a signal-to-noise ratio guided noise adaptive network, characterized in that: The following steps are involved: Step 1: construct a low-light image enhancement system based on a signal-to-noise ratio guided noise adaptive network, wherein the low-light image enhancement system is provided with an image acquisition module and a signal-to-noise ratio guided noise adaptive network SNA-Net, wherein the signal-to-noise ratio guided noise adaptive network SNA-Net is provided with an input layer, wherein the output end of the input layer is connected to the input end of a CNN encoder, an SNA encoder and a signal-to-noise ratio map calculation module, wherein the output ends of the CNN encoder, the SNA encoder and the signal-to-noise ratio map calculation module are all connected to the input end of a feature fusion module, wherein the output end of the CNN encoder is also connected to the input end of a CNN decoder via a convolution residual block, and the input end of the CNN decoder is also connected to the output end of the feature fusion module; Step 2: The image acquisition module acquires a low-light image I and passes it to the signal-to-noise ratio guided noise adaptive network SNA-Net; Step 3: The input layer of the signal-to-noise ratio guided noise adaptive network SNA-Net obtains the low-light image I and passes it to the CNN encoder, SNA encoder and signal-to-noise ratio map calculation module; Step 4: The signal-to-noise ratio map calculation module calculates the low-light image I to obtain a corresponding signal-to-noise ratio map S, and then passes the signal-to-noise ratio map S to the feature fusion module and the SNA encoder; Step 5: The CNN encoder performs short-distance encoding on the low-light image I to capture local information of the low-light image I and obtain short-distance feature data a i , and passed to the feature fusion module and the convolution residual block; The SNA encoder uses the signal-to-noise ratio map S as prior knowledge to calculate the self-attention score, and then performs long-distance encoding on the low-light image I to filter out noise and redundant information to obtain long-distance feature data b i , and passed to the feature fusion module; Step 6: The convolution residual block performs convolution on the short-distance feature data a i Perform convolution residual operation to obtain convolution residual data d, and pass it to the CNN decoder; The feature fusion module combines the short-distance feature data a i and long-distance feature data b i Perform feature fusion operation to obtain fused feature data c i , and passed to the CNN decoder; Step 7: The CNN decoder processes the convolution residual data d and the fusion feature data c i Perform decoding operation to obtain enhanced image 2. The low-light image enhancement method based on the signal-to-noise ratio guided noise adaptive network according to claim 1, characterized in that: In the step 1, the CNN encoder is provided with a first CNN encoding block, a second CNN encoding block, a third CNN encoding block and a fourth CNN encoding block connected in sequence, the SNA encoder is provided with a first SNA block, a second SNA block, a third SNA block and a fourth SNA block connected in sequence, the feature fusion module is provided with a first feature fusion block, a second feature fusion block, a third feature fusion block and a fourth feature fusion block, and the CNN decoder is provided with a first CNN decoding block, a second CNN decoding block, a third CNN decoding block and a fourth CNN decoding block connected in sequence; The output ends of the first CNN encoding block and the first SNA block are connected to the input end of the first feature fusion block, and the output end of the first feature fusion block is connected to the input end of the fourth CNN decoding block; The output ends of the second CNN encoding block and the second SNA block are connected to the input end of the second feature fusion block, and the output end of the second feature fusion block is connected to the input end of the third CNN decoding block; The output ends of the third CNN encoding block and the third SNA block are connected to the input end of the third feature fusion block, and the output end of the third feature fusion block is connected to the input end of the second CNN decoding block; The output ends of the fourth CNN encoding block and the fourth SNA block are connected to the input end of the fourth feature fusion block, and the output end of the fourth feature fusion block is connected to the input end of the first CNN decoding block.

3. The low-light image enhancement method based on a signal-to-noise ratio guided noise adaptive network according to claim 1, characterized in that: In step 4, firstly, the grayscale image I corresponding to the low-light image I is calculated. g , and then calculate the signal-to-noise ratio map S according to the following formula, S∈R H×W ; Among them, Denoise() means using a mean filter for denoising, abs means absolute value, N∈R H×W Estimated noise map.

4. The low-light image enhancement method based on a signal-to-noise ratio guided noise adaptive network according to claim 2, characterized in that: The first CNN coding block, the second CNN coding block, the third CNN coding block and the fourth CNN coding block have the same structure, and are all provided with a 3×3 convolution layer and a downsampling layer, and the output end of the 3×3 convolution layer is built-in with a LeakyReLU activation function.

5. The low-light image enhancement method based on a signal-to-noise ratio guided noise adaptive network according to claim 2, characterized in that: The first SNA block, the second SNA block, the third SNA block, and the fourth SNA block have the same structure, and are all provided with a noise adaptive self-attention module NASA and a dual domain refinement feedforward network DRFN; The noise adaptive self-attention module NASA in each SNA block uses the signal-to-noise ratio map S as prior knowledge to calculate the self-attention score Attn of the corresponding SNA block, and then uses the self-attention score Attn to perform long-distance encoding on the low-light image I or the output data of the previous SNA block to obtain the self-attention feature of the current SNA block, and then further refines the self-attention feature through the dual-domain refinement feedforward network DRFN to obtain the long-distance feature data b of the current SNA block. i .

6. The low-light image enhancement method based on a signal-to-noise ratio guided noise adaptive network according to claim 5, characterized in that: The noise adaptive self-attention module NASA obtains the self-attention features of the current SNA block through the following steps: Step A1: The first normalization layer in the noise adaptive self-attention module NASA obtains the output data of the low-light image I or the previous SNA block, generates tensor data, and then passes it to the query matrix encoding unit, the key matrix K encoding unit, the value matrix V encoding unit and the first addition unit; Step A2: the query matrix encoding unit encodes the tensor data through a first 1×1 convolution layer and a first 3×3 depth convolution layer to generate a query matrix Q, and passes it to a first multiplication unit; The key matrix K encoding unit encodes the tensor data through a second 1×1 convolution layer and a second 3×3 depth convolution layer to generate a key matrix K, and performs matrix transposition on the key matrix K before passing it to the first multiplication unit; The value matrix V encoding unit encodes the tensor data through the third 1×1 convolution layer and the third 3×3 depth convolution layer to generate a value matrix V, and passes it to the second multiplication unit; The expressions for generating the query matrix Q, key matrix K, and value matrix V are as follows: Among them, W d represents a 1×1 point-by-point convolution, W p Represents a 3×3 depth convolution, X represents tensor data; Step A3: The first multiplication unit multiplies the query matrix Q and the transposed key matrix K T Multiply them together to get the first multiplication matrix, and pass it to the sparse self-attention branch and the dense self-attention branch; Step A4: the dense self-attention branch calculates the first attention data through the first Softmax function according to the first multiplication matrix, and passes it to the second addition unit; The sparse self-attention branch uses the normalized signal-to-noise ratio map S′ as a mask to filter out the areas in the first multiplication matrix whose signal-to-noise ratio values ​​are lower than the signal-to-noise ratio threshold, and then calculates the score of the areas whose signal-to-noise ratio values ​​are higher than or equal to the signal-to-noise ratio threshold through the second Softmax function to obtain the second attention data, and passes it to the second addition unit; Step A5: The second addition unit weights the first attention data with the first attention weight w1, and weights the second attention data with the second attention weight w2, and adds the weighted first attention score and the second attention score, and then multiplies them with the value matrix V through the second multiplication unit to obtain the self-attention score Attn, which is expressed as follows: Attn=w1·D attn +w2·S attn D attn =Softmax(QK T / α)V S attn =Softmax(QK T / α+(1-S′)σ)V Where α is a scaling parameter that controls the size of the dot product between the query matrix Q and the key matrix K; σ is a negative scalar -1e9, D attn is the dense self-attention score, S attn is the sparse self-attention score; Step A6: The second multiplication unit passes the self-attention score Attn to the fourth 1×1 convolutional layer for convolution operation, and then adds the convolution result to the tensor data element by element through the first addition unit to obtain the self-attention feature of the current SNA block, and passes it to the dual-domain refinement feedforward network DRFN.

7. The low-light image enhancement method based on a signal-to-noise ratio guided noise adaptive network according to claim 4, characterized in that: The dual domain refinement feedforward network DRFN further refines the self-attention features through the following steps: Step B1: The second normalization layer in the dual domain refinement feedforward network DRFN normalizes the self-attention feature, and then the fifth 1×1 convolution layer performs a convolution operation on the normalized feature to obtain the fifth convolution data, and passes it to the sixth 1×1 convolution layer, the fourth 3×3 depth convolution layer and the fast Fourier transform unit FFT; The expression for the fifth convolution data is as follows: Y′=W d (LN(Y)) Among them, Y is the self-attention feature, LN is the layer normalization, Y′ is the fifth convolution data, and W d Represents a 1×1 point-by-point convolution; Step B2: the sixth 1×1 convolution layer performs a convolution operation on the fifth convolution data, and then performs a depth convolution operation on the fifth 3×3 depth convolution layer to obtain fifth depth convolution data, and pass it to the first element-by-element multiplication unit; The expression for the fifth depth convolution data is as follows: Y″=W p W d (Y′) Among them, W p represents a 3×3 depth convolution, and Y″ represents the fifth depth convolution data; The fourth 3×3 depth convolution layer performs a depth convolution operation on the fifth convolution data, and then activates it through a GELU activation function to obtain first activation data, and passes it to the element-by-element multiplication unit; The expression for getting the first activation data is as follows: in, is the GELU activation function, and Y″′ represents the first activation data; Step B3: the fast Fourier transform unit FFT performs fast Fourier transform on the fifth convolution data to obtain FFT data, and then multiplies the FFT data element by element with the weight w of the dual-domain refinement feedforward network DRFN through the second element-by-element multiplication unit, and then performs inverse fast Fourier transform through the inverse fast Fourier transform unit IFFT, and then activates through the GEGLU activation function to obtain second activation data, and passes it to the connection unit; The expression for obtaining the second activation data is as follows: in, is the fast Fourier transform, is the inverse fast Fourier transform, w is the weight of the dual domain refinement feedforward network DRFN, Y f is the second activation data, ζ() represents the GEGLU function; The first element-by-element multiplication unit performs element-by-element multiplication on the fifth depth convolution data and the first activation data to obtain first multiplication data, and passes the first multiplication data to the connection unit; The expression for the first multiplication data is as follows: AND s =Y″⊙Y″′ Among them, ⊙ represents element-by-element multiplication, Y s is the first multiplication data; Step B4: the connection unit performs a connection operation on the second activation data and the first multiplication data, and then performs a convolution operation on the seventh 1×1 convolution layer to obtain seventh convolution data, and passes it to the third addition unit; Step B5: The third adding unit adds the self-attention feature and the seventh convolution data element by element to obtain the long-distance feature data b of the current SNA block. i ; Get the long-distance characteristic data b of the current SNA block i The expression is as follows: in, Represents a join operation.

8. The low-light image enhancement method based on a signal-to-noise ratio guided noise adaptive network according to claim 2, characterized in that: The first feature fusion block, the second feature fusion block, the third feature fusion block and the fourth feature fusion block have the same structure, and are all provided with a third multiplication unit, a fourth multiplication unit, a fourth addition unit, a global average pooling layer, a first fully connected layer, a second fully connected layer, a fifth multiplication unit, a sixth multiplication unit, and a fifth addition unit, and a Sigmoid function is built in the output end of the second fully connected layer; In step 6, the feature fusion block obtains the corresponding short-range feature data a i and long-distance feature data b i And perform feature fusion operation to obtain the fusion feature data c of the current feature fusion block i , including the following steps: Step C1: The third multiplication unit in the current feature fusion block uses the normalized signal-to-noise ratio map as a mask and combines it with the short-range feature data a i multiplying to obtain third multiplication data, and passing it to the fourth addition unit and the fifth multiplication unit; The fourth multiplication unit in the current feature fusion block combines the normalized signal-to-noise ratio map with the long-distance feature data b i multiplying to obtain fourth multiplication data, and passing the fourth addition unit and the sixth multiplication unit; Step C2: The fourth addition unit adds the third multiplication data and the fourth multiplication data element by element, and then performs a global average pooling operation through the global average pooling layer, and then sequentially performs a full connection operation through the first fully connected layer and the second fully connected layer, and then obtains a short-distance feature fusion weight w after normalization by the Sigmoid function. a And the long-distance feature fusion weight w a , and the short-distance feature fusion weight w a Passed to the fifth multiplication unit, the long-distance feature fusion weight w d Passed to the sixth multiplication unit; Step C3: The fifth multiplication unit is fused with weight w through short-distance features a weighting the third multiplication data to obtain fifth multiplication data, and transmitting the fifth multiplication data to a fifth adding unit; The sixth multiplication unit is fused with weights w through long-distance features. d weighting the fourth multiplication data to obtain sixth multiplication data, and transmitting the sixth multiplication data to a fifth adding unit; Step C4: the fifth adding unit adds the fifth multiplication data and the sixth multiplication data element by element to obtain the feature fusion data c of the current feature fusion block. i , the expression is as follows: c i =SGFF(a i ×(1-S′)+b i ×S′) Among them, SGFF represents the feature fusion operation, S′ represents the normalized signal-to-noise ratio map, and 1-S′ represents using the normalized signal-to-noise ratio map as a mask.

Citation Information

Patent Citations

  • Image enhancement method, vehicle snapshot method, device and medium

    CN115984133A

  • Double-branch fusion low-illumination image enhancement method

    CN116977208A

  • HSV and Transform-based unsupervised low-illumination image enhancement method

    CN118247192A

  • Infrared image dynamic range adaptive enhancement method and system based on signal-to-noise ratio perception

    CN118396914A

  • Low-illumination image enhancement method based on light effect perception

    CN118761945A