Low-light image enhancement method based on illumination grouping and mask attention
By designing a multi-scale guided group attention module and a center mask attention-aware feature reconstruction module, the problem of local detail information and noise processing in low-light image enhancement is solved, achieving high-quality image brightening and denoising effects.
Patent Information
- Application Number
- CN202510044777.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Existing low-light image enhancement methods are insufficient in terms of local detail information extraction and noise processing, especially the Mamba method, which lacks effective means for local dark light brightening and noise removal.
We design a multi-scale guided group attention module and a feature reconstruction module based on center mask attention perception. We recover local detail information through multi-scale feature learning, multi-branch multi-receptive field convolution and group attention mechanism, and process noise through pixel rearrangement, center mask convolution and mask attention operation.
It effectively restores local details in low-light images, improves image contrast, and effectively removes noise during the brightening process, thus improving image quality.
Smart Images

Figure CN119941601B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of image and video processing and computer vision, and specifically relates to a low-light image enhancement method based on illumination grouping and mask attention. Background Technology
[0002] The goal of low-light image enhancement methods is to make low-light images increasingly approximate normal-light images. Compared to normal-light images, low-light images lack local detail and normal illumination information, and also suffer from color distortion and noise issues after brightening. Therefore, low-light images pose certain challenges to related downstream tasks, including pedestrian and road detection under low-light conditions, autonomous driving, and security monitoring.
[0003] Solving problems in low-light images typically involves both hardware and software aspects. On the hardware side, using a flash or extending the exposure time can compensate for the lack of image brightness, but this may result in an overly dark background and blurring of moving pedestrians or objects. Alternatively, image brightening software can increase image brightness, but this can lead to loss of detail and noise.
[0004] Traditional low-light image enhancement methods are mainly divided into histogram equalization methods and methods based on Retinex theory. For histogram equalization, Kim et al. (Y. Kim, Contrast enhancement using brightness preserving bi-histogram equalization, IEEE transactions on Consumer Electronics, vol. 43, no. 1, pp. 1-8, 1997.) proposed mean-preserving bi-histogram equalization to retain image brightness information. Abdullah-Al-Wadud et al. (M. Abdullah-Al-Wadud, M. Hasanul Kabir, M. Ali Akber Dewan, and O Chae, A dynamic histogram equalization for image contrast enhancement, IEEE transactions on consumer electronics, vol. 53, no. 2, pp. 593-600, 2007.) used histogram segmentation and gray-level redistribution methods to preserve image details. In recent years, some methods have attempted to combine histogram equalization with deep learning methods. For example, Zhang et al. (F. Zhang, Y. Shao, Y. Sun, K. Zhu, C. Gao, and N. Sang, Unsupervised low-light image enhancement via histogram equalization prior, arXiv preprint arXiv, 2112.01766, 2021.) proposed combining histogram equalization and convolution-based methods to enhance the brightness features of images. Regarding Retinex-based methods, Fu et al. (X. Fu, Y. Sun, M. Li Wang, Y. Huang, X. Zhang, X. Ding, Anovel retinex based approach for image enhancement with illumination adjustment, IEEE International Conference on Acoustics, Speech and Signal Processing, 2014, 1190-1194.) utilized a logarithmic-free Retinex method to preserve edge information of image features.In recent years, there have been attempts to combine Retinex methods with deep learning methods. For example, Cai et al. (Y.Cai, H.Bian, J.Lin, H.Wang, R.Timofte, Y.Zhang, Retinexformer: One-Stage retinex-based Transformer for low-light image enhancement, in Proceedings of the IEEE International Conference on ComputerVision, 2023, 12504-12513.) proposed a single-stage framework combining Retinex theory and Transformer for image brightening and degradation information recovery operations.
[0005] In recent years, an increasing number of works in the field of low-light image enhancement have adopted deep learning-based methods, which have achieved better performance compared to traditional low-light image enhancement methods. For example, Xu et al. (X.Xu, R.Wang, C.Fu, J.Jia, SNR-aware low-light image enhancement, in Proceedings of the Computer Vision and Pattern Recognition, 2022, 17714-17724.) proposed combining convolutional methods and signal-to-noise ratio-aware Transformers to learn local and non-local features of images respectively. In addition, there have been some recent popular works, such as using Mamba-based methods to solve the global dimming problem. Specifically, the Mamba method can uncover the relationships between global features of an image through multi-directional global scanning operations. However, the Mamba-based method still has shortcomings in terms of local detail information discovery and local dimming enhancement. To address this issue, Weng et al. (J.Weng, Z.Yan, Y.Tai, J.Qian, J.Yang, J.Li, MambaLLIE: Implicit Retinex-Aware Low Light Enhancement with Global-then-Local State Space, arXiv preprint arXiv, 2405.16105, 2024.) proposed IRSK to separate positive and negative lighting information, but this method relies on prior lighting knowledge. Some methods combine Mamba with CNN-based methods. For example, Zhang et al. (X.Zhang,H.Zeng,J.Pan,Q.Shen,Y.Chen,LLEMamba:Low-Light Enhancement via Relighting-Guided Mamba with Deep Unfolding Network,arXiv preprint arXiv,2406.01028,2024.) chose to perform convolutional operations first, followed by Mamba operations. However, this simple combination cannot fully utilize the feature discovery capabilities of both Mamba and CNN.Zou et al. (W. Zou, H. Gao, W. Yang, T. Liu, Wave-Mamba: Wavelet State Space Model for Ultra-High-Definition Low-Light Image Enhancement, in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, 1534-1543.) transferred features from each stage in the encoder to the corresponding stage in the decoder to enhance the feature representation within the decoder; however, this lacks the discovery of multi-scale local features within the decoder stages. Li et al. (G. Li, K. Zhang, T. Wang, M. Li, B. Zhao, X. Li, Semi-LLIE: Semi-supervised Contrastive Learning with Mamba-based Low-light Image Enhancement, arXiv preprint arXiv, 2409.16604, 2024.) used two convolutional operations with different kernel sizes to discover local multi-scale features; however, this method lacks feature discovery across multiple receptive fields.
[0006] Because low-light image enhancement involves noise, some methods attempt to improve the noise level by employing different denoising operations. For example, Yi et al. (X.Yi,H.Xu,H.Zhang,L.Tang,J.Ma,Diff-retinex:Rethinking low-light image enhancement with a generative diffusion model,in Proceedings of the Conference on ComputerVision,2023,12302-12311.) proposed adding Gaussian noise to simulate a noisy environment and training the model to remove the simulated noise. However, Gaussian noise cannot completely simulate the noise distribution in real images. Makwana et al. (D.Makwana, G.Deshmukh, O.Susladkar, S.Mittal, S.Chandra, LIVENet: A novel network for real-world low-light image denoising and enhancement, in Proceedings of the Winter Conference on Applications of Computer Vision, 2024, 5856-5865.) achieved denoising by using low-rank matrices to preserve the original image structure; however, low-rank matrix representations can lead to the loss of some detailed information. Summary of the Invention
[0007] The purpose of this invention is to overcome the problems existing in the background technology and provide a low-light image enhancement method based on illumination grouping and mask attention. This method designs a multi-scale guided grouping attention module to discover local detail information and solve the problem of local dark light, and designs a feature reconstruction module based on center mask attention to reconstruct and remove noise in the brightening process.
[0008] To achieve the above objectives, the technical solution of the present invention is: a low-light image enhancement method based on illumination grouping and mask attention, comprising:
[0009] Step A: Preprocess the input image, including image data pairing, data cropping, and image enhancement, to obtain the training dataset;
[0010] Step B: Design a low-light image enhancement network based on illumination grouping and mask attention, including an image brightening module, a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction, and a feature output module;
[0011] Step C: Design a loss function for updating the network parameters of the low-light image enhancement network based on illumination grouping and masked attention in step B;
[0012] Step D: Train the low-light image enhancement network based on illumination grouping and mask attention from step B using the training dataset to obtain the trained low-light image enhancement network based on illumination grouping and mask attention.
[0013] Step E: Use the trained low-light image enhancement network based on illumination grouping and mask attention obtained in Step D to test the image and obtain the predicted normal illumination image.
[0014] In one embodiment of the present invention, step A is specifically implemented as follows:
[0015] Step A1: Pair the low-light image with the label image;
[0016] Step A2: Randomly crop each low-light image and the label image in the same way to obtain an image of size H×W×3, where H and W are the height and width of the cropped image;
[0017] Step A3: Randomly apply one of the following 8 data augmentation methods to the paired low-light image and label image: keep the original image, flip vertically, rotate 90 degrees counterclockwise, rotate 90 degrees counterclockwise and flip vertically, rotate 180 degrees counterclockwise, rotate 180 degrees counterclockwise and flip vertically, rotate 270 degrees counterclockwise, rotate 270 degrees counterclockwise and flip vertically.
[0018] In one embodiment of the present invention, step B is specifically implemented as follows:
[0019] Step B1: Design an image brightening module, including 2D convolution and depthwise separable convolution, to brighten low-light images. Generate preliminary brightened image features Where H and W represent the height and width of the image feature, and C represents the channel of the image feature;
[0020] Step B2: Design a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction. The overall architecture includes an encoder, a decoder, and an intermediate deep feature processing module. The degradation recovery module includes S stages, used to degrade and recover the preliminary brightened image features obtained in step B1 into features.
[0021] Step B3: Design a feature output module, including a two-dimensional convolution operation and a residual connection operation, to generate a normal illumination image from the recovered features y' obtained in step B2.
[0022] In one embodiment of the present invention, step B1 is specifically implemented as follows:
[0023] Step B11: To obtain the brightening features, the low-light images in the training dataset obtained in Step A are... Averaging is performed along the channel dimension to obtain the prior lighting features. Then the low-light image I input With prior features of illumination I p The features are concatenated along the channel dimension, and the merged feature size is H×W×4. Then, the brightening features are obtained through two-dimensional convolution and depthwise separable convolution operations. The specific implementation method is as follows:
[0024] Z lu =DWConv 5×5 (Conv 1×1 (Cat(I input ,I p )))
[0025] Where Cat(·) represents the merge operation, Conv 1×1 (·) represents a two-dimensional convolution operation with a 1×1 kernel, DWConv 5×5 (·) indicates a depthwise separable convolution operation with a 5×5 kernel;
[0026] Step B12, Brightening Feature Z lu After a two-dimensional convolution operation, and compared with the original image I input Perform multiplication and residual join operations to generate the brightened image. The specific formula is as follows:
[0027]
[0028] in, Conv represents element-wise multiplication. 1×1 (·) represents a two-dimensional convolution operation with a 1×1 kernel;
[0029] Step B13: The brightened image I generated in step B12... lu After two-dimensional convolution, feature T is obtained. s The specific formula is as follows:
[0030] T s =Conv 3×3 (I lu )
[0031] Among them, Conv 3×3 (·) represents a two-dimensional convolution operation with a 3×3 kernel.
[0032] In one embodiment of the present invention, step B2 is specifically implemented as follows:
[0033] Step B21: Design a multi-scale guided group attention module. The module first performs multi-scale feature learning operation, then performs multi-branch multi-receptive field convolution operation, and finally performs group attention operation to restore local details and brighten local dark light features.
[0034] Step B22: Design an encoder, including two stages. s=0 and Stage s=1 Stage s=0 In sequence, it consists of an attention mechanism for illumination fusion, a two-dimensional selective scanning operation, and a multi-scale guided grouped attention module. s=1 It consists of an illumination fusion attention mechanism and a two-dimensional selective scan operation. The illumination fusion attention mechanism (J.Bai, Y.Yin, Q.He, Y.Li, X.Zhang, RetinexMamba: Retinex-based Mamba for low-light image enhancement, arXiv preprint arXiv, 2405.03349, 2024.) introduces illumination features for cross-attention operations, allowing the model to better focus on dark areas that need enhancement. The two-dimensional selective scan operation (S.Wang, Q.Tao, and Z.Tang, Resvmunetx: A low-light enhancement network based on vmamba, arXiv preprint arXiv, 2401.10166, 2024.) includes a scan expansion operation, an S6 block, and a scan merging operation, which solves the image dark light problem from a global perspective. The S6 block represents the interaction between each feature in the sequence and the features of the previous scan. By using a compressed hidden state, the quadratic complexity is reduced to linear complexity.
[0035] Step B23: Design a feature reconstruction module with center mask attention perception. The module first performs pixel rearrangement operation, then center mask convolution operation, and finally mask attention operation to learn how to reconstruct simulated noise features using limited non-noise information.
[0036] Step B24: Design an intermediate deep feature processing module, namely Stage s=2 It includes an attention mechanism for illumination fusion, a two-dimensional selection scanning operation, and a feature reconstruction module for center mask attention perception.
[0037] Step B25: Design a decoder, including two stages. s=3 and Stage s=4 Stage s=3 Stage consists of an attention mechanism for illumination fusion and a two-dimensional selective scanning operation. s=4 In sequence, it consists of an attention mechanism for illumination fusion, a two-dimensional selection scanning operation, and a multi-scale guided group attention module.
[0038] In one embodiment of the present invention, step B21 is specifically implemented as follows:
[0039] Step B211: Design a multi-scale feature learning operation, including a three-branch depthwise separable convolution operation, with different kernel sizes used in each of the three branches, to achieve the discovery of multi-scale local features. The specific formula is as follows:
[0040]
[0041] in, This represents the features obtained after illumination fusion attention mechanism and 2D selection scanning, where s represents the s-th stage, LP(·) represents the linear mapping operation, σ represents the GELU activation function, and DWConv c×c This indicates a depthwise separable convolution operation with a kernel size of c×c, where c takes the value of 1, 3, or 5. ∑ represents the summation of the output features of the three depthwise separable convolutions.
[0042] Step B212: Design multi-branch, multi-receptive-field convolution operations to process the features obtained in step B211. Perform four-branch convolution operations;
[0043] First, a convolutional block with kernel e×f consists of a two-dimensional convolution operation with kernel e×f, a batch normalization operation, and a ReLU activation function. The convolutional layer in the first branch consists of a convolutional block with kernel 1×1. The convolutional layers in the second, third, and fourth branches are all composed of convolutional blocks with kernel 1×1, convolutional blocks with kernel 1×n, convolutional blocks with kernel n×1, and dilated convolutional blocks with kernel n×n, combined in sequence. In the second, third, and fourth branches, n takes the values 3, 5, and 7, respectively.
[0044] After the input feature P undergoes a multi-branch, multi-receptive-field convolution operation, it yields four branches of output features. These output features are then concatenated along the channel dimension to obtain the final output feature. Next, a 2D convolution operation with a 3×3 kernel is used to reduce the channel dimension of X to match the size of the input feature P. Finally, a residual connection operation is performed with the input feature to obtain the output feature. The specific process is as follows:
[0045] Y = Conv 3×3 (X)+Conv 1×1 (P)
[0046] Among them, Conv 3×3 (·) and Conv 1×1 These represent 3×3 convolution and 1×1 convolution operations, respectively.
[0047] Step B213: Design a group attention operation, including a group attention mechanism and a gated recurrent unit; first, implement the group attention mechanism, which distinguishes between bright and dark features from a global perspective by grouping the bright and dark features within a single image, and brightens the correctly grouped local dark features; first, randomly initialize learnable clustering feature embeddings. Where S represents the number of groups within a single image, C represents the channel size, and then E is assigned to the initial clustering features. The clustering feature originates from randomly initialized grouping features used to discover bright and dark areas within the image, which are used to generate the first query feature. Here, t represents the number of times the learnable clustering feature is updated. A total of three grouping attention mechanisms and gated recurrent units are required. After the t-th attention mechanism and gated recurrent unit, the output feature is processed using layer normalization and linear mapping to obtain a new query feature, which is then passed to the next (t+1)-th grouping attention mechanism and gated recurrent unit operation. After the final grouping attention mechanism and gated recurrent unit operation, a spatial location embedding operation propagates the grouping information of bright and dark areas within the image to the image features to obtain the output.
[0048] Specifically, firstly, position embeddings are added to the multi-scale feature Y generated in step B212, and then it is flattened into a 2D feature. Where N = H × W, then Y' obtains feature A through a multilayer perceptron and layer normalization, and A generates a key matrix through layer normalization and linear mapping. Sum matrix At the same time y t The query matrix is generated through layer normalization and linear mapping. The specific formula is as follows:
[0049] Q t =LN(y t W q K = LN(A)W kV=LN(A)W v
[0050] Where LN(·) represents layer normalization, W q W k W v This represents a linear mapping operation;
[0051] When performing the attention mechanism for the t-th time, the query matrix Q is first performed. t The dot product operation of the bond matrix K is used to learn the relationship between different positions within a feature, thus obtaining the attention weights. Then, a weighted average is used to stabilize the attention weights; finally, to reference the features of important positions, the weighted average attention weights are multiplied by the value matrix V to obtain the output features. The specific operation process is as follows:
[0052]
[0053]
[0054] in, This indicates regularization, the superscript T of K indicates the transpose operation, and Softmax(·) indicates the normalization exponential function. This indicates a weighted average operation;
[0055] Design gated recurrent units to update the learnable cluster feature representation; specifically, utilize the current cluster feature y t and feature O t Intermediate features are generated through a gated loop unit. The specific formula is expressed as follows:
[0056] G t =GRU(O t ,y t )
[0057] Wherein, GRU(·) represents a gated loop unit operation;
[0058] Finally, layer normalization and multilayer perceptron are used to generate the (t+1)th updated clustering feature. The formula is expressed as follows:
[0059] y t+1 =MLP(LN(G t ))+G t
[0060] Where LN(·) represents layer normalization and MLP represents multilayer perceptron operation.
[0061] In one embodiment of the present invention, step B23 is specifically implemented as follows:
[0062] Step B231: Design a pixel rearrangement operation by analyzing image features. All pixels within the pixel array are rearranged to disrupt the correlation between local noise, making the noise distribution more random and uniform. Feature U is obtained after the pixel rearrangement operation. Specifically, the feature is first obtained through a reshaping operation. Then, the features are obtained by rearranging them. Where r represents the scaling factor, and finally C×r 2 The size is The matrix is arranged from left to right and from top to bottom and then reshaped to obtain the output features.
[0063] The above rearrangement process shows that the local relationships of noise are disrupted, making the local noise distribution randomized after rearrangement, which makes the subsequent feature reconstruction more robust when facing different noises.
[0064] Step B232: Design a center mask convolution operation. Utilize limited features within a local area to reconstruct features from the center mask, increasing the model's ability to reconstruct and remove local noise features. First, perform a masking operation on the center features within the local area; specifically, using the features at each pixel position... Divide the matrix into a 3×3 shape around the center, and perform a masking operation on the center pixel, i.e., assign a value of 0 to it and assign a value of 1 to the other pixels. The specific operation is as follows:
[0065]
[0066] Where the variables m∈[i-1,i+1] and n∈[j-1,j+1], when dividing the matrix into 3×3 with the edge pixel features as the center, pixel filling operation is performed on the missing pixels at the edges. * denotes the dot product operation. This represents a 3×3 center mask matrix centered at pixels i and j.
[0067] Then, H×W u” in the image i,j The center mask matrix is first subjected to a 3×3 convolution operation, followed by two 1×1 convolution operations to reconstruct the features of the center pixel of the mask and obtain the output feature B.
[0068] Step B233: Design a window-based multi-head masking attention mechanism. First, perform random masking within the global feature range of the image. Specifically, first, randomly set the masking ratio p, which is randomly generated between 0.6 and 0.9. Then, randomly select pixels to be masked, ensuring that the ratio of the number of pixels to be masked to the total number of pixels in the image is p. Randomly generate a sequence L of pixel positions to be masked within the image. If a pixel belongs to sequence L, perform the masking operation; otherwise, retain the original feature value. After the random masking operation, obtain the output features. The specific formula is expressed as follows:
[0069]
[0070] Where (m,n) represents the pixel position with height m and width n in the image;
[0071] Subsequently, the features are divided into h heads along the channel dimension for multi-head attention, and the query matrix is obtained through linear mapping. Key matrix Sum matrix Output features are obtained through multi-head attention. The specific formula is expressed as follows:
[0072]
[0073] in, R is used for regularization operations, and R is used for position embedding operations.
[0074] Finally, after the reshaping operation, the size of feature A becomes Furthermore, the multi-head features are spliced together to obtain...
[0075] Step B234: Design a masked attention operation to further enhance the model's ability to remove severe noise by increasing the range and quantity of noise. Unlike the center masked convolution operation in step B232, the masked attention operation in this step increases the quantity and range of simulated noise, enabling the model to better reconstruct and remove noise features when facing severe noise in a global scope. First, the output features of step B232 are... Perform window-based partitioning to obtain the partitioned features. Where Z represents the side length of the partitioned window; multiple window features are processed using a window-based multi-head masking attention mechanism, followed by window merging to obtain intermediate features. Finally, the output features are obtained by performing layer normalization and multilayer perceptron operations, and adding residual connections. The specific formula is expressed as follows:
[0076] Y = B + WM(MWA(LN(B')))
[0077] M = Y + MLP(LN(Y))
[0078] Where LN(·) represents the layer normalization operation, MWA(·) represents the window-based multi-head mask attention mechanism, MLP(·) represents the multilayer perceptron operation, and WM(·) represents the window merging operation. Specifically, the shape is first obtained through the first reshaping operation. The features are then reshaped in a second operation to reduce the Z-axis in the feature dimension. 2 It is split into Z×Z, and finally, after a third reshaping operation, the output feature with shape H×W×C is obtained.
[0079] In one embodiment of the present invention, step C is specifically implemented as follows:
[0080] The network parameters are updated using the L1 norm loss, as expressed by the following formula:
[0081]
[0082] Among them, y i For the true value, The predicted value from the network output, where N represents the number of samples.
[0083] In one embodiment of the present invention, step D is specifically implemented as follows:
[0084] The processed training dataset obtained in step A is divided into J batches, with each batch containing Z pairs of images. For the z-th low-light image I in the j-th batch, the enhanced image I is obtained by using the low-light image enhancement network based on illumination grouping and mask attention in step B. output The loss of the enhanced image is calculated using the loss function designed in step C to update the network parameters. The Adam optimizer is used to update the network parameters, resulting in a trained low-light image model based on illumination grouping and mask attention.
[0085] The present invention also provides a computer-readable storage medium having stored thereon computer program instructions that can be executed by a processor, wherein when the processor executes the computer program instructions, it can implement the steps of the method described above.
[0086] Compared to existing technologies, this invention offers the following advantages: First, to address the poor local detail learning and local bright / dark grouping issues inherent in the Mamba method, this invention designs a multi-scale guided grouping attention module. This module includes multi-scale feature learning operations, multi-branch multi-receptive-field convolution operations, and a grouping attention mechanism. Multi-scale feature learning and multi-branch multi-receptive-field convolution operations address image detail loss from the perspective of local multi-receptive fields, while the grouping attention mechanism enhances local brightness and restores image contrast. To address the noise problem during brightening, a center mask attention-based perceptual feature reconstruction module is designed. This module includes pixel rearrangement operations, center mask convolution operations, and mask attention operations. Pixel rearrangement operations disrupt the correlation between local noises, making the noise distribution more random and uniform. Center mask convolution operations train the model to reconstruct features from the center mask using useful information within a local area. Mask attention operations enhance the model's noise removal capabilities when facing global noise. Attached Figure Description
[0087] Figure 1 This is a flowchart illustrating the implementation of the method of the present invention.
[0088] Figure 2 This is a structural diagram of a low-light image enhancement network based on illumination grouping and mask attention in an embodiment of the present invention.
[0089] Figure 3 This is a structural diagram of the image brightening module in an embodiment of the present invention.
[0090] Figure 4 This is a structural diagram of the group attention module based on multi-scale guidance in an embodiment of the present invention.
[0091] Figure 5 This is a structural diagram of the center mask attention-aware feature reconstruction module in an embodiment of the present invention. Detailed Implementation
[0092] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.
[0093] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0094] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0095] This invention provides a low-light image enhancement method based on illumination grouping and mask attention, comprising:
[0096] Step A: Preprocess the input image, including image data pairing, data cropping, and image enhancement, to obtain the training dataset;
[0097] Step B: Design a low-light image enhancement network based on illumination grouping and mask attention, including an image brightening module, a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction, and a feature output module;
[0098] Step C: Design a loss function for updating the network parameters of the low-light image enhancement network based on illumination grouping and masked attention in step B;
[0099] Step D: Train the low-light image enhancement network based on illumination grouping and mask attention from step B using the training dataset to obtain the trained low-light image enhancement network based on illumination grouping and mask attention.
[0100] Step E: Use the trained low-light image enhancement network based on illumination grouping and mask attention obtained in Step D to test the image and obtain the predicted normal illumination image.
[0101] The following is a detailed implementation process of the present invention.
[0102] This invention provides a low-light image enhancement method based on illumination grouping and mask attention, the implementation flowchart of which is shown below. Figure 1 As shown, the network structure diagram is as follows: Figure 2 As shown, it includes the following steps:
[0103] Step A: Preprocess the input image, including image data pairing, data cropping, and image enhancement, to obtain the training dataset;
[0104] Step B: Design a low-light image enhancement network based on illumination grouping and mask attention. The network includes an image brightening module, a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction, and a feature output module.
[0105] Step C: Design the loss function for updating the network parameters in step B;
[0106] Step D: Use the processed training data from Step A to train the low-light image enhancement network from Step B, and obtain a trained low-light image enhancement network based on illumination grouping and mask attention.
[0107] Step E: Use the trained low-light image enhancement network obtained in step D to test the image and obtain the predicted normal-light image.
[0108] Further, step A includes the following steps:
[0109] Step A1: Pair the low-light image with the label image;
[0110] Step A2: Randomly crop each low-light image and the label image in the same way to obtain an image of size H×W×3, where H and W are the height and width of the cropped image;
[0111] Step A3: Randomly apply one of the following 8 data augmentation methods to the paired low-light image and label image: keep the original image, flip vertically, rotate 90 degrees counterclockwise, rotate 90 degrees counterclockwise and flip vertically, rotate 180 degrees counterclockwise, rotate 180 degrees counterclockwise and flip vertically, rotate 270 degrees counterclockwise, rotate 270 degrees counterclockwise and flip vertically.
[0112] Further, step B includes the following steps:
[0113] Step B1: Design an image brightening module, including 2D convolution and depthwise separable convolution, to brighten low-light images. Generate preliminary brightened image features
[0114] Step B2: Design a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction. The overall architecture includes an encoder, a decoder, and an intermediate deep feature processing module. This module contains S stages, used to degrade and recover the preliminary brightened image features obtained in Step B1 into features.
[0115] Step B3: Design a feature output module, including a two-dimensional convolution operation and a residual connection operation, to generate a normal illumination image from the recovered features y' obtained in step B2.
[0116] Furthermore, such as Figure 3 As shown, step B1 includes the following steps:
[0117] Step B11: To obtain the brightening features, the low-light image obtained in step A is... Averaging is performed along the channel dimension to obtain the prior lighting features. Where H and W represent the height and width of the image features. Then, the low-light image I... input With prior features of illumination I p The features are concatenated along the channel dimension, resulting in a combined feature size of H×W×4. Brightening features are then obtained through 2D convolution and depthwise separable convolution operations. Where C represents the channel size of the feature, and the specific implementation is as follows:
[0118] Z lu =DWConv 5×5 (Conv 1×1 (Cat(I input ,I p )))
[0119] Where Cat(·) represents the merge operation, Conv 1×1 (·) represents a two-dimensional convolution operation with a 1×1 kernel, DWConv 5×5 (·) indicates a depthwise separable convolution operation with a 5×5 kernel;
[0120] Step B12, Brightening Feature Z lu After a two-dimensional convolution operation, and compared with the original image I input Perform multiplication and residual join operations to generate the brightened image. The specific formula is as follows:
[0121]
[0122] in, Conv represents element-wise multiplication. 1×1 (·) represents a two-dimensional convolution operation with a 1×1 kernel;
[0123] Step B13: The brightened image I generated in step B12... lu After two-dimensional convolution, feature T is obtained. s The specific formula is as follows:
[0124] T s =Conv 3×3 (I lu )
[0125] Among them, Conv 3×3 (·) represents a two-dimensional convolution operation with a 3×3 kernel;
[0126] Further, step B2 includes the following steps:
[0127] Step B21: Design a multi-scale guided group attention module. The module first performs multi-scale feature learning operation, then performs multi-branch multi-receptive field convolution operation, and finally performs group attention operation to restore local details and brighten local dark light features.
[0128] Step B22: Design an encoder that includes two stages. s=0 and Stage s=1 Stage s=0 In sequence, it consists of an attention mechanism for illumination fusion, a two-dimensional selective scanning operation, and a multi-scale guided grouped attention module. s=1 It consists of an illumination fusion attention mechanism and a two-dimensional selective scan operation. The illumination fusion attention mechanism used (J. Bai, Y. Yin, Q. He, Y. Li, X. Zhang, RetinexMamba: Retinex-based Mamba for low-light image enhancement, arXiv preprint arXiv, 2405.03349, 2024.) introduces illumination features for cross-attention operations, allowing the model to better focus on dark areas that need enhancement. The two-dimensional selective scan operation used (S. Wang, Q. Tao, and Z. Tang, Resvmunetx: A low-light enhancement network based on vmamba, arXiv preprint arXiv, 2401.10166, 2024.) includes a scan expansion operation, an S6 block, and a scan merging operation, which addresses the low-light problem of the image from a global perspective. The S6 block represents the interaction between each feature in the sequence and the features of the previous scan, using a compressed hidden state to reduce the quadratic complexity to linear complexity.
[0129] Step B23: Design a feature reconstruction module with center mask attention perception. The module first performs a pixel rearrangement operation, then a center mask convolution operation, and finally a mask attention operation to learn how to reconstruct simulated noise features using limited non-noise information.
[0130] Step B24: Design an intermediate deep feature processing module, namely Stage s=2 It includes an attention mechanism for illumination fusion, a two-dimensional selection scanning operation, and a feature reconstruction module for center mask attention perception.
[0131] Step B25: Design a decoder comprising two stages. s=3 and Stage s=4 Stages=3 It consists of an attention mechanism for illumination fusion and a two-dimensional selective scanning operation. Stage s=4 In sequence, it consists of an attention mechanism for illumination fusion, a two-dimensional selection scanning operation, and a multi-scale guided group attention module.
[0132] Furthermore, such as Figure 4 As shown, step B21 includes the following steps:
[0133] Step B211: Design a multi-scale feature learning operation. This operation mainly includes a three-branch depthwise separable convolution operation, with different kernel sizes used in each branch to achieve multi-scale local feature discovery. The specific formula is as follows:
[0134]
[0135] in, The term represents the features obtained after illumination fusion attention mechanism and 2D selection scanning, where s represents the s-th stage in the method. LP(·) represents the linear mapping operation, σ represents the GELU activation function, and DWConv c×c This indicates a depthwise separable convolution operation with a kernel size of c×c, where c takes values of 1, 3, and 5. '∑' indicates the summation of the output features of the three depthwise separable convolutions.
[0136] Step B212: Design multi-branch, multi-receptive-field convolution operations. Apply the features obtained in step B211 to... Perform four-branch convolution operations.
[0137] First, a convolutional block with kernel size e×f consists of a two-dimensional convolution operation with kernel size e×f, a batch normalization operation, and a ReLU activation function. The convolutional layer in the first branch consists of a convolutional block with kernel size 1×1. The convolutional layers in the second, third, and fourth branches are all composed of convolutional blocks with kernel size 1×1, convolutional blocks with kernel size 1×n, convolutional blocks with kernel size n×1, and dilated convolutional blocks with kernel size n×n, combined in sequence. In the second, third, and fourth branches, n takes the values 3, 5, and 7, respectively.
[0138] After the input feature P undergoes a multi-branch, multi-receptive-field convolution operation, it yields four branches of output features. These output features are then concatenated along the channel dimension to obtain the final output feature. Next, a 2D convolution operation with a 3×3 kernel is used to reduce the channel dimension of X to match the size of the input feature P. Finally, a residual connection operation is performed with the input feature to obtain the output feature. The specific process is as follows:
[0139] Y = Conv 3×3 (X)+Conv1×1 (P)
[0140] Among them, Conv 3×3 (·) and Conv 1×1 These represent 3×3 convolution and 1×1 convolution operations, respectively.
[0141] Step B213: Design a group attention operation. This mainly includes a group attention mechanism and a gated recurrent unit. First, the group attention mechanism is implemented. This mechanism groups the bright and dark features within a single image to better distinguish between bright and dark features from a global perspective, and brightens the correctly grouped local dark features. First, learnable clustering feature embeddings are randomly initialized. Where S represents the number of groups within a single image, and C represents the channel size. Then, E is assigned to the initial clustering features. The clustering features are derived from randomly initialized grouping features used to discover bright and dark areas within the image, and are used to generate the first query feature. Here, t represents the number of times the learnable clustering features are updated. This method requires three grouping attention mechanisms and gated recurrent units. After the t-th attention mechanism and gated recurrent unit operation, the output features are processed using layer normalization and linear mapping to obtain new query features, which are then passed to the next t+1-th grouping attention mechanism and gated recurrent unit operation. After the final grouping attention mechanism and gated recurrent unit operation, a spatial location embedding operation propagates the grouping information of bright and dark areas within the image to the image features to obtain the output.
[0142] Specifically, firstly, position embeddings are added to the multi-scale feature Y generated in step B212, and then it is flattened into a 2D feature. Where N = H × W. Subsequently, Y' obtains features A through a multilayer perceptron and layer normalization, while A generates a key matrix through layer normalization and linear mapping. Sum matrix At the same time y t The query matrix is generated through layer normalization and linear mapping. The specific formula is as follows:
[0143] Q t =LN(y t W q K = LN(A)W k V=LN(A)W v
[0144] Where LN(·) represents layer normalization, W q W k W v This represents a linear mapping operation.
[0145] When performing the attention mechanism for the t-th time, the query matrix Q is first performed. t The dot product operation of the bond matrix K is used to learn the relationship between different positions within a feature, thus obtaining the attention weights. Then, a weighted average is used to stabilize the attention weights. Finally, to reference the features of important positions, the weighted average attention weights are multiplied by the value matrix V to obtain the output features. The specific operation process is as follows:
[0146]
[0147]
[0148] in, The expression represents regularization, the superscript 'T' of K indicates the rank transformation operation, and Softmax(·) represents the normalized exponential function. This indicates a weighted average operation.
[0149] Design gated recurrent units to update the learnable cluster feature representation. Specifically, utilize the current cluster feature y t and feature O t Intermediate features are generated through a gated loop unit. The specific formula is expressed as follows:
[0150] G t =GRU(O t ,y t )
[0151] GRU(·) represents a gated cyclic unit operation.
[0152] Finally, layer normalization and multilayer perceptron are used to generate the (t+1)th updated clustering feature. The formula is expressed as follows:
[0153] y t+1 =MLP(LN(G t ))+G t
[0154] Where LN(·) represents layer normalization and MLP represents multilayer perceptron operation.
[0155] Furthermore, such as Figure 5 As shown, step B23 includes the following steps:
[0156] Step B231: Design a pixel rearrangement operation by analyzing image features. All pixels within the pixel array are rearranged to disrupt the correlation between local noise, making the noise distribution more random and uniform. Feature U is obtained after this pixel rearrangement. Specifically, the feature is first obtained through a reshaping operation. Then, the features are obtained by rearranging them. Where r represents the scaling factor, and finally 'C×r' 2 'A size of The matrix is arranged from left to right and from top to bottom and then reshaped to obtain the output features.
[0157] The above rearrangement process shows that the local relationships of noise are disrupted, making the local noise distribution randomized after rearrangement, which makes the subsequent feature reconstruction more robust when facing different noises.
[0158] Step B232: Design the center mask convolution operation. This involves using finite features within a local area to reconstruct features from the center mask, increasing the model's ability to reconstruct and remove local noise features. First, a masking operation is performed on the center features within the local area. Specifically, this involves using the features at each pixel position... Divide the matrix into a 3×3 shape around the center, and perform a masking operation on the center pixel, i.e., assign it a value of '0', and assign the value of '1' to the other pixels. The specific operation is as follows:
[0159]
[0160] Where the variables m∈[i-1,i+1] and n∈[j-1,j+1], when dividing the matrix into 3×3 with the edge pixel features as the center, pixel filling operation is performed on the missing pixels at the edges, and '*' represents the dot product operation. This represents a 3×3 center mask matrix centered at pixels i and j.
[0161] Next, the 'H×W' u' values within the image were analyzed. i,j The center mask matrix is first subjected to a 3×3 convolution operation, followed by two 1×1 convolution operations to reconstruct the features of the center pixel of the mask and obtain the output feature B.
[0162] Step B233: Design a window-based multi-head masking attention mechanism. First, perform random masking within the global feature range of the image. Specifically, first, randomly set the masking ratio p, which is randomly generated between 0.6 and 0.9. Then, randomly select pixels to be masked, ensuring that the ratio of the number of pixels to be masked to the total number of pixels in the image is p. Randomly generate a sequence L of pixel positions to be masked within the image. If a pixel belongs to sequence L, the masking operation is performed; otherwise, the pixel retains its original feature value. After the random masking operation, the output features are obtained. The specific formula is expressed as follows:
[0163]
[0164] Where (m,n) represents the pixel position with height m and width n in the image.
[0165] Subsequently, the features are divided into h heads along the channel dimension for multi-head attention, and the query matrix is obtained through linear mapping. Key matrix Sum matrix Output features are obtained through multi-head attention. The specific formula is expressed as follows:
[0166]
[0167] in, R is used for regularization operations, and R is used for position embedding operations.
[0168] Finally, after the reshaping operation, the size of feature A becomes Furthermore, the multi-head features are spliced together to obtain...
[0169] Step B234: Design a masked attention operation to further enhance the model's ability to remove severe noise by increasing the range and quantity of noise. Unlike the center masked convolution operation in step B232, the masked attention operation in this step increases the quantity and range of simulated noise, enabling the model to better reconstruct and remove noise features when facing severe noise globally. First, the output features from step B232... Perform window-based partitioning to obtain the partitioned features. Where Z represents the side length of the partitioned window. Multiple window features are processed using a window-based multi-head masking attention mechanism, followed by window merging to obtain intermediate features. Finally, the output features are obtained by performing layer normalization and multilayer perceptron operations, and adding residual connections. The specific formula is expressed as follows:
[0170] Y = B + WM(MWA(LN(B')))
[0171] M = Y + MLP(LN(Y))
[0172] Where LN(·) represents the layer normalization operation, MWA(·) represents the window-based multi-head mask attention mechanism, MLP(·) represents the multilayer perceptron operation, and WM(·) represents the window merging operation. Specifically, the first reshaping operation first obtains a shape of... The features are then reshaped in a second reshaping operation to remove the 'Z' in the feature dimension. 2 It is split into 'Z×Z', and finally, after a third reshaping operation, the output feature with the shape 'H×W×C' is obtained.
[0173] Further, step C includes the following steps:
[0174] Step C: Use the L1 norm loss to update the network parameters. The specific formula is as follows:
[0175]
[0176] Among them, y i For the true value, The predicted value from the network output, where N represents the number of samples.
[0177] Furthermore, step D includes the following steps:
[0178] Step D: The training dataset, after being processed in Step A, is divided into J batches, each containing Z pairs of images. For the z-th low-light image I in the j-th batch, the enhanced image I is obtained by using the low-light image enhancement network based on illumination grouping and mask attention in Step B. output The loss of the enhanced image is calculated using the loss function designed in step C to update the network parameters. The Adam optimizer is used to update the network parameters, resulting in a trained low-light image model based on illumination grouping and mask attention.
[0179] The present invention also provides a computer-readable storage medium having stored thereon computer program instructions that can be executed by a processor, wherein when the processor executes the computer program instructions, it can implement the steps of the method described above.
[0180] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0181] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0182] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0183] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0184] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A low-light image enhancement method based on illumination grouping and masked attention, characterized in that, include: Step A: Preprocess the input image, including image data pairing, data cropping, and image enhancement, to obtain the training dataset; Step B: Design a low-light image enhancement network based on illumination grouping and mask attention, including an image brightening module, a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction, and a feature output module; Step C: Design a loss function for updating the network parameters of the low-light image enhancement network based on illumination grouping and masked attention in step B; Step D: Train the low-light image enhancement network based on illumination grouping and mask attention from step B using the training dataset to obtain the trained low-light image enhancement network based on illumination grouping and mask attention. Step E: Use the trained low-light image enhancement network based on illumination grouping and mask attention obtained in step D to test the image and obtain the predicted normal illumination image. The specific implementation steps of step B are as follows: Step B1: Design an image brightening module, including 2D convolution and depthwise separable convolution, to brighten low-light images. Generate preliminary brightened image features Where H and W represent the height and width of the image feature, and C represents the channel of the image feature; Step B2: Design a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction. The overall architecture includes an encoder, a decoder, and an intermediate deep feature processing module. The degradation recovery module includes S stages, used to degrade and recover the preliminary brightened image features obtained in step B1 into features. Step B3: Design a feature output module, including a two-dimensional convolution operation and a residual connection operation, to generate a normal illumination image from the recovered features y' obtained in step B2. The specific implementation steps of step B2 are as follows: Step B21: Design a multi-scale guided group attention module. The module first performs multi-scale feature learning operation, then performs multi-branch multi-receptive field convolution operation, and finally performs group attention operation to restore local details and brighten local dark light features. The group attention operation includes a group attention mechanism and a gated recurrent unit. Step B22: Design an encoder, including two stages. s=0 and Stage s=1 Stage s=0 In sequence, it consists of an attention mechanism for illumination fusion, a two-dimensional selective scanning operation, and a multi-scale guided grouped attention module. s=1 It consists of an attention mechanism for illumination fusion and a two-dimensional selective scanning operation; The illumination fusion attention mechanism used introduces illumination features for cross-attention operations, allowing the model to better focus on dark areas that need enhancement. The two-dimensional selection scan operation used includes scan expansion, S6 block and scan merging operations, which solve the image dark light problem from a global perspective. The S6 block represents the interaction between each feature in the sequence and the features of the previous scan. By using a compressed hidden state, the quadratic complexity is reduced to linear complexity. Step B23: Design a feature reconstruction module with center mask attention perception. The module first performs pixel rearrangement operation, then center mask convolution operation, and finally mask attention operation to learn how to reconstruct simulated noise features using limited non-noise information. Step B24: Design an intermediate deep feature processing module, namely Stage s=2 This includes the attention mechanism for illumination fusion, the two-dimensional selection scanning operation, and the feature reconstruction module for center mask attention perception; Step B25: Design a decoder, including two stages. s=3 and Stage s=4 Stage s=3 Stage consists of an attention mechanism for illumination fusion and a two-dimensional selective scanning operation. s=4 In sequence, it consists of an attention mechanism for illumination fusion, a two-dimensional selection scanning operation, and a multi-scale guided group attention module.
2. The low-light image enhancement method based on illumination grouping and mask attention according to claim 1, characterized in that, The specific implementation steps of step A are as follows: Step A1: Pair the low-light image with the label image; Step A2: Randomly crop each low-light image and the label image in the same way to obtain an image of size H×W×3, where H and W are the height and width of the cropped image; Step A3: Randomly apply one of the following 8 data augmentation methods to the paired low-light image and label image: keep the original image, flip vertically, rotate 90 degrees counterclockwise, rotate 90 degrees counterclockwise and flip vertically, rotate 180 degrees counterclockwise, rotate 180 degrees counterclockwise and flip vertically, rotate 270 degrees counterclockwise, rotate 270 degrees counterclockwise and flip vertically.
3. The low-light image enhancement method based on illumination grouping and mask attention according to claim 1, characterized in that, The specific implementation steps of step B1 are as follows: Step B11: To obtain the brightening features, the low-light images in the training dataset obtained in Step A are... Averaging is performed along the channel dimension to obtain the prior lighting features. Then the low-light image I input With prior features of illumination I p The features are concatenated along the channel dimension, and the merged feature size is H×W×4. Then, the brightening features are obtained through two-dimensional convolution and depthwise separable convolution operations. The specific implementation method is as follows: Z lu =DWConv 5×5 (Conv 1×1 (Cat(I input ,I p ))) Where Cat(·) represents the merge operation, Conv 1×1 (·) represents a two-dimensional convolution operation with a 1×1 kernel, DWConv 5×5 (·) indicates a depthwise separable convolution operation with a 5×5 kernel; Step B12, Brightening Feature Z lu After a two-dimensional convolution operation, and compared with the original image I input Perform multiplication and residual join operations to generate the brightened image. The specific formula is as follows: in, Conv represents element-wise multiplication. 1×1 (·) represents a two-dimensional convolution operation with a 1×1 kernel; Step B13: The brightened image I generated in step B12... lu After two-dimensional convolution, feature T is obtained. s The specific formula is as follows: T s =Conv 3×3 (I lu ) Among them, Conv 3×3 (·) represents a two-dimensional convolution operation with a 3×3 kernel.
4. The low-light image enhancement method based on illumination grouping and mask attention according to claim 1, characterized in that, The specific implementation steps of step B21 are as follows: Step B211: Design a multi-scale feature learning operation, including a three-branch depthwise separable convolution operation, with different kernel sizes used in each of the three branches, to achieve the discovery of multi-scale local features. The specific formula is as follows: in, This represents the features obtained after illumination fusion attention mechanism and 2D selection scanning, where s represents the s-th stage, LP(·) represents the linear mapping operation, σ represents the GELU activation function, and DWConv C×C This represents a depthwise separable convolution operation with a kernel size of C×c, where C takes the value of 1, 3, or 5. ∑ represents the summation of the output features of the three depthwise separable convolutions. Step B212: Design multi-branch, multi-receptive-field convolution operations to process the features obtained in step B211. Perform four-branch convolution operations; First, a convolutional block with kernel e×f consists of a two-dimensional convolution operation with kernel e×f, a batch normalization operation, and a ReLU activation function. The convolutional layer in the first branch consists of a convolutional block with kernel 1×1. The convolutional layers in the second, third, and fourth branches are all composed of convolutional blocks with kernel 1×1, convolutional blocks with kernel 1×n, convolutional blocks with kernel n×1, and dilated convolutional blocks with kernel n×n, combined in sequence. In the second, third, and fourth branches, n takes the values 3, 5, and 7, respectively. After the input feature P undergoes a multi-branch, multi-receptive-field convolution operation, it yields four branches of output features. These output features are then concatenated along the channel dimension to obtain the final output feature. Next, a 2D convolution operation with a 3×3 kernel is used to reduce the channel dimension of X to match the size of the input feature P. Finally, a residual connection operation is performed with the input feature to obtain the output feature. The specific process is as follows: Y=Conv 3×3 (X)+Conv 1×1 (P) Among them, Conv 3×3 (·) and Conv 1×1 These represent 3×3 convolution and 1×1 convolution operations, respectively. Step B213: Design a group attention operation, including a group attention mechanism and a gated recurrent unit; first, implement the group attention mechanism, which distinguishes between bright and dark features from a global perspective by grouping the bright and dark features within a single image, and brightens the correctly grouped local dark features; first, randomly initialize learnable clustering feature embeddings. Where S represents the number of groups within a single image, C represents the channel size, and then E is assigned to the initial clustering features. The clustering feature originates from randomly initialized grouping features used to discover bright and dark areas within the image, which are used to generate the first query feature. Here, t represents the number of times the learnable clustering feature is updated. A total of three grouping attention mechanisms and gated recurrent units are required. After the t-th attention mechanism and gated recurrent unit, the output feature is processed using layer normalization and linear mapping to obtain a new query feature, which is then passed to the next (t+1)-th grouping attention mechanism and gated recurrent unit operation. After the final grouping attention mechanism and gated recurrent unit operation, a spatial location embedding operation propagates the grouping information of bright and dark areas within the image to the image features to obtain the output. Specifically, firstly, position embeddings are added to the multi-scale feature Y generated in step B212, and then it is flattened into a 2D feature. Where N = H × W, then Y' obtains feature A through a multilayer perceptron and layer normalization, and A generates a key matrix through layer normalization and linear mapping. Sum matrix At the same time y t The query matrix is generated through layer normalization and linear mapping. The specific formula is as follows: Q t =LN(y t )W q ,K=LN(A)W k ,V=LN(A)W v Where LN(·) represents layer normalization, W q W k W v This represents a linear mapping operation; When performing the attention mechanism for the t-th time, the query matrix Q is first performed. t The dot product operation of the bond matrix K is used to learn the relationship between different positions within a feature, thus obtaining the attention weights. Then, a weighted average is used to stabilize the attention weights; finally, to reference the features of important positions, the weighted average attention weights are multiplied by the value matrix V to obtain the output features. The specific operation process is as follows: in, This indicates regularization, the superscript T of K indicates the transpose operation, and Softmax(·) indicates the normalization exponential function. This indicates a weighted average operation; Design gated recurrent units to update the learnable cluster feature representation; specifically, utilize the current cluster feature y t and feature O t Intermediate features are generated through a gated loop unit. The specific formula is expressed as follows: G t =GRU(O t ,y t ) Wherein, GRU(·) represents the gated loop unit operation; Finally, layer normalization and multilayer perceptron are used to generate the (t+1)th updated clustering feature. The formula is expressed as follows: y t+1 =MLP(LN(G t ))+G t Where LN(·) represents layer normalization and MLP represents multilayer perceptron operation.
5. The low-light image enhancement method based on illumination grouping and mask attention according to claim 1, characterized in that, The specific implementation steps of step B23 are as follows: Step B231: Design a pixel rearrangement operation by analyzing image features. All pixels within the pixel array are rearranged to disrupt the correlation between local noise, making the noise distribution more random and uniform. Feature U is obtained after the pixel rearrangement operation. Specifically, the feature is first obtained through a reshaping operation. Then, the features are obtained by rearranging them. Where r represents the scaling factor, and finally C×r 2 The size is The matrix is arranged from left to right and from top to bottom and then reshaped to obtain the output features. Step B232: Design a center mask convolution operation. Utilize limited features within a local area to reconstruct features from the center mask, increasing the model's ability to reconstruct and remove local noise features. First, perform a masking operation on the center features within the local area; specifically, using the features at each pixel position... Divide the matrix into a 3×3 shape around the center, and perform a masking operation on the center pixel, i.e., assign a value of 0 to it and assign a value of 1 to the other pixels. The specific operation is as follows: Where the variables m∈[i-1,i+1] and n∈[j-1,j+1], when dividing the matrix into 3×3 with the edge pixel features as the center, pixel filling operation is performed on the missing pixels at the edges. * denotes the dot product operation. This represents a 3×3 center mask matrix centered at pixels i and j. Then, H×W u” in the image i,j The center mask matrix is first subjected to a 3×3 convolution operation, followed by two 1×1 convolution operations to reconstruct the features of the center pixel of the mask and obtain the output feature B. Step B233: Design a window-based multi-head masking attention mechanism. First, perform random masking within the global feature range of the image. Specifically, first, randomly set the masking ratio p, which is randomly generated between 0.6 and 0.
9. Then, randomly select pixels to be masked, ensuring that the ratio of the number of pixels to be masked to the total number of pixels in the image is p. Randomly generate a sequence L of pixel positions to be masked within the image. If a pixel belongs to sequence L, perform the masking operation; otherwise, retain the original feature value. After the random masking operation, obtain the output features. The specific formula is expressed as follows: Where (m,n) represents the pixel position with height m and width n in the image; Subsequently, the features are divided into h heads along the channel dimension for multi-head attention, and the query matrix is obtained through linear mapping. Key matrix Sum matrix Output features are obtained through multi-head attention. The specific formula is expressed as follows: in, R is used for regularization operations, and R is used for position embedding operations. Finally, after the reshaping operation, the size of feature A becomes Furthermore, the multi-head features are spliced together to obtain... Step B234: Design a masked attention operation to further enhance the model's ability to remove severe noise by increasing the range and quantity of noise. Unlike the center masked convolution operation in step B232, the masked attention operation in this step increases the quantity and range of simulated noise, enabling the model to better reconstruct and remove noise features when facing severe noise in a global scope. First, the output features of step B232 are... Perform window-based partitioning to obtain the partitioned features. Where Z represents the side length of the partitioned window; multiple window features are processed using a window-based multi-head masking attention mechanism, followed by window merging to obtain intermediate features. Finally, the output features are obtained by performing layer normalization and multilayer perceptron operations, and adding residual connections. The specific formula is expressed as follows: Y = B + WM(MWA(LN(B'))) M = Y + MLP(LN(Y)) Where LN(·) represents the layer normalization operation, MWA(·) represents the window-based multi-head mask attention mechanism, MLP(·) represents the multilayer perceptron operation, and WM(·) represents the window merging operation. Specifically, the shape is first obtained through the first reshaping operation. The features are then reshaped in a second operation to reduce the Z-axis in the feature dimension. 2 It is split into Z×Z, and finally, after a third reshaping operation, the output feature with shape H×W×C is obtained.
6. The low-light image enhancement method based on illumination grouping and mask attention according to claim 1, characterized in that, The specific implementation method of step C is as follows: The network parameters are updated using the L1 norm loss, as expressed by the following formula: Among them, y i For the true value, The predicted value from the network output, where N represents the number of samples.
7. The low-light image enhancement method based on illumination grouping and mask attention according to claim 1, characterized in that, Step D is implemented as follows: The processed training dataset obtained in step A is divided into J batches, with each batch containing Z pairs of images. For the z-th low-light image I in the j-th batch, the enhanced image I is obtained by using the low-light image enhancement network based on illumination grouping and mask attention in step B. output ; The loss of the enhanced image is calculated using the loss function designed in step C, which is then used to update the network parameters. The Adam optimizer is used to update the network parameters, resulting in a trained low-light image model based on illumination grouping and mask attention.
8. A computer-readable storage medium having stored thereon computer program instructions executable by a processor, wherein when the processor executes the computer program instructions, it is able to implement the steps of the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Low-illumination image enhancement method based on Retinex and deep learning
CN111968044A
Low-illumination image enhancement method, system and device and medium
CN114359073A