Low-illumination image enhancement method based on illumination grouping and mask attention

By introducing a multi-scale guided grouping attention module and a feature reconstruction module based on central mask attention perception in the low-illumination image enhancement method, local details information discovery and noise problems are solved, and a better low-illumination image brightening effect is achieved.

CN119941601AActive Publication Date: 2025-05-06FUZHOU UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510044777.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-06
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

The existing low-illumination image enhancement methods have shortcomings in the exploration of local details and local dark light brightening, and the noise problem during the brightening process has not been effectively solved.

Method used

A low-illumination image enhancement method based on illumination packet and mask attention is designed, and a multi-scale guided grouping attention module and a feature reconstruction module based on central mask attention perception are used to discover local detailed information and solve noise problems, respectively.

Benefits of technology

It effectively restores local details of low-illumination images, brightens local dark light areas, and effectively removes noise during the brightening process, improving the quality of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941601A_ABST
    Figure CN119941601A_ABST
Patent Text Reader

Abstract

The invention relates to a low-illumination image enhancement method based on illumination grouping and mask attention, and belongs to the field of image and video processing and computer vision. The method comprises the following steps: carrying out operations such as preprocessing, image data pairing, data cutting and image enhancement on an input image to obtain a training data set; designing a low-illumination image enhancement network based on illumination grouping and mask attention, wherein the network comprises an image brightening module, a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction, and a feature output module; designing a loss function for network parameter updating; training a low-illumination image enhancement network based on illumination grouping and mask attention to obtain a trained low-illumination image enhancement network; and testing the trained network by using the test set to obtain a predicted normal illumination image. According to the method, the low-light image can be brightened, and the problems of local detail missing, local dark light and noise after image brightening are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of image and video processing and computer vision, and in particular relates to a low-illumination image enhancement method based on illumination grouping and mask attention. Background Art

[0002] The purpose of low-light image enhancement methods is to make low-light images approach normal bright-light images. Compared with normal bright-light images, low-light images lack local detail information and normal light information, and have color distortion and noise problems after brightening. Therefore, low-light images bring certain challenges to related downstream tasks, including pedestrian detection and road detection under low-light conditions, autonomous driving, and safety monitoring.

[0003] Solving the problems caused by low-light images usually involves both hardware and software. In terms of hardware, using a flash or extending the exposure time to compensate for the lack of image brightness may result in a dark background and blurring of moving pedestrians or objects. In addition, the brightness of the image can be increased by using image brightening software, but this may result in a lack of image details and noise.

[0004] Traditional low-light image enhancement methods are mainly divided into histogram equalization methods and methods based on Retinex theory. For the histogram equalization method, Kim et al. (Y. Kim, Contrast enhancement using brightness preserving bi-histogram equalization, IEEE transactions on Consumer Electronics, vol. 43, no. 1, pp. 1-8, 1997.) proposed mean-preserving bi-histogram equalization to retain the brightness information of the image. Abdullah-Al-Wadud et al. (M. Abdullah-Al-Wadud, M. Hasanul Kabir, M. Ali Akber Dewan, and O Chae, A dynamic histogram equalization for image contrast enhancement, IEEE transactions on consumer electronics, vol. 53, no. 2, pp. 593-600, 2007.) use histogram segmentation and grayscale redistribution methods to retain image details. In recent years, some methods have attempted to combine histogram equalization with deep learning methods. For example, Zhang et al. (F. Zhang, Y. Shao, Y. Sun, K. Zhu, C. Gao, and N. Sang, Unsupervised low-light image enhancement via histogram equalization prior, arXiv preprint arXiv, 2112.01766, 2021.) proposed combining histogram equalization and convolution-based methods to increase the bright light features of the image. For methods based on Retinex theory, Fu et al. (X. Fu, Y. Sun, M. Li Wang, Y. Huang, X. Zhang, X. Ding, A novel retinex based approach for image enhancement with illumination adjustment, IEEE International Conference on Acoustics, Speech and Signal Processing, 2014, 1190-1194.) used the Retinex method without logarithmic transformation to retain the edge information of image features.In recent years, there have been attempts to combine the Retinex method with deep learning methods. For example, Cai et al. (Y. Cai, H. Bian, J. Lin, H. Wang, R. Timofte, Y. Zhang, Retinexformer: One-Stage retinex-based Transformer for low-light image enhancement, in Proceedings of the IEEE International Conference on Computer Vision, 2023, 12504-12513.) proposed a single-stage framework combining Retinex theory and Transformer to perform image brightening and restoration of degradation information.

[0005] In recent years, in the field of low-light image enhancement, more and more works have used deep learning-based methods. Compared with traditional low-light image enhancement methods, deep learning-based methods have achieved better performance. For example, Xu et al. (X.Xu, R.Wang, C.Fu, J.Jia, SNR-aware low-light image enhancement, in Proceedings of the Computer Vision and Pattern Recognition, 2022, 17714-17724.) proposed a combination of convolutional methods and signal-to-noise ratio-aware Transformer to learn local and non-local features of images respectively. In addition, there are some popular works recently, such as using Mamba-based methods to solve global dark light problems. Specifically, the Mamba method can discover the relationship between global features of the image through multi-directional global scanning operations. However, the Mamba-based method still lacks in local detail information discovery and local dark light brightening. In order to solve this problem, Weng et al. (J. Weng, Z. Yan, Y. Tai, J. Qian, J. Yang, J. Li, MambaLLIE: Implicit Retinex-Aware Low Light Enhancement with Global-then-Local State Space, arXiv preprint arXiv, 2405.16105, 2024.) proposed IRSK to separate positive illumination information and negative illumination information, but this method needs to rely on prior knowledge of illumination. Some methods combine Mamba and CNN-based methods. For example, Zhang et al. (X. Zhang, H. Zeng, J. Pan, Q. Shen, Y. Chen, LLE Mamba: Low-Light Enhancement via Relighting-Guided Mamba with Deep Unfolding Network, arXiv preprint arXiv, 2406.01028, 2024.) chose to perform convolution operation first and then Mamba operation. However, the simple combination method cannot fully utilize the ability of both Mamba and CNN to discover features.Zou et al. (W.Zou, H.Gao, W.Yang, T.Liu, Wave-Mamba: Wavelet State Space Model for Ultra-High-Definition Low-Light Image Enhancement, in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, 1534-1543.) transfer the features of each stage in the encoder to the corresponding stage in the decoder to enhance the feature representation in the decoder, but this lacks the exploration of multi-scale local features within the decoder stage. Li et al. (G.Li, K.Zhang, T.Wang, M.Li, B.Zhao, X.Li, Semi-LLIE: Semi-supervised Contrastive Learning with Mamba-based Low-light Image Enhancement, arXiv preprint arXiv, 2409.16604, 2024.) use two convolution operations with different convolution kernel sizes to explore local multi-scale features, but this method lacks feature exploration of multiple receptive fields.

[0006] Since there will be noise information in the image during the low-light image brightening process, some methods try to use different denoising operations to improve the noise information in the image. For example, Yi et al. (X.Yi, H.Xu, H.Zhang, L.Tang, J.Ma, Diff-retinex: Rethinking low-light image enhancement with a generative diffusion model, in Proceedings of the Conference on Computer Vision, 2023, 12302-12311.) proposed adding Gaussian noise to simulate the noise environment and train the model to remove the simulated noise, but Gaussian noise cannot completely simulate the noise distribution in the real image. Makwana et al. (D.Makwana, G.Deshmukh, O.Susladkar, S.Mittal, S.Chandra, LIVENet: A novel network for real-world low-light image denoising and enhancement, in Proceedings of the Winter Conference on Applications of Computer Vision, 2024, 5856-5865.) use a low-rank matrix to retain the original image structure to achieve denoising, but the low-rank matrix representation will bring about the problem of missing some detail information. Summary of the invention

[0007] The purpose of the present invention is to overcome the problems existing in the background technology and provide a low-light image enhancement method based on illumination grouping and mask attention. The method designs a multi-scale guided grouping attention module to explore local detail information and solve the local dark light problem, and designs a feature reconstruction module based on center mask attention perception to reconstruct and remove the noise existing in the brightening process.

[0008] To achieve the above object, the technical solution of the present invention is: a low-illumination image enhancement method based on illumination grouping and mask attention, comprising:

[0009] Step A: preprocessing the input image, including operations of image data pairing, data cropping and image enhancement, to obtain a training data set;

[0010] Step B, designing a low-light image enhancement network based on illumination grouping and mask attention, including an image brightening module, a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction, and a feature output module;

[0011] Step C, designing a loss function for updating network parameters of the low-light image enhancement network based on illumination grouping and mask attention in step B;

[0012] Step D: using the training data set to train the low-illumination image enhancement network based on illumination grouping and mask attention in step B to obtain a trained low-illumination image enhancement network based on illumination grouping and mask attention;

[0013] Step E: Use the trained low-illumination image enhancement network based on illumination grouping and mask attention obtained in step D to test the image and obtain a predicted normal illumination image.

[0014] In one embodiment of the present invention, the specific implementation steps of step A are as follows:

[0015] Step A1, pairing the low-light image with the label image;

[0016] Step A2: randomly crop each low-light image and the label image to obtain an image of size H×W×3, where H and W are the height and width of the cropped image;

[0017] Step A3: Randomly use one of the following eight data augmentation methods on the paired low-light image and the label image: keep the original image, flip upside down, rotate 90 degrees counterclockwise, rotate 90 degrees counterclockwise and flip upside down, rotate 180 degrees counterclockwise, rotate 180 degrees counterclockwise and flip upside down, rotate 270 degrees counterclockwise, and rotate 270 degrees counterclockwise and flip upside down.

[0018] In one embodiment of the present invention, the specific implementation steps of step B are as follows:

[0019] Step B1: Design an image brightening module, including two-dimensional convolution and depthwise separable convolution, to brighten low-light images. Generate preliminary brightened image features Where H and W represent the height and width of the image feature, and C represents the channel of the image feature;

[0020] Step B2: Design a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction. The overall architecture includes an encoder, a decoder, and an intermediate deep feature processing module. The degradation recovery module includes S stages, which are used to restore the initial brightened image features obtained in step B1 to feature

[0021] Step B3: Design a feature output module, including a two-dimensional convolution operation and a residual connection operation, to generate a normal illumination image from the restored feature y' obtained in step B2

[0022] In one embodiment of the present invention, the specific implementation steps of step B1 are as follows:

[0023] Step B11: To obtain brightening features, the low-light image in the training data set obtained in step A is Perform an averaging operation in the channel dimension to obtain the lighting prior features Then the low-light image I input With the lighting prior feature I p After concatenation in the channel dimension, the size of the merged feature is H×W×4, and then the brightening feature is obtained through two-dimensional convolution operation and depth-separable convolution operation. The specific implementation method is:

[0024] Z lu =DWConv 5×5 (Conv 1×1 (Cat(I input ,I p )))

[0025] Among them, Cat(·) represents the merging operation, Conv 1×1 (·) indicates a two-dimensional convolution operation with a convolution kernel of 1×1, DWConv 5×5 (·) indicates a depthwise separable convolution operation with a kernel of 5×5;

[0026] Step B12: Brighten feature Z lu After two-dimensional convolution operation, and with the original image I input Perform multiplication and residual connection operations to generate a brightened image The specific formula is as follows:

[0027]

[0028] in, Represents element-wise multiplication, Conv 1×1 (·) indicates a two-dimensional convolution operation with a convolution kernel of 1×1;

[0029] Step B13: The brightened image I generated in step B12 is lu After two-dimensional convolution, we get feature T s , the specific formula is as follows:

[0030] T s =Conv 3×3 (I lu )

[0031] Among them, Conv 3×3 (·) represents a two-dimensional convolution operation with a convolution kernel of 3×3.

[0032] In one embodiment of the present invention, the specific implementation steps of step B2 are as follows:

[0033] Step B21, design a multi-scale guided group attention module, which first undergoes a multi-scale feature learning operation, then a multi-branch multi-receptive field convolution operation, and finally a group attention operation, for local detail recovery and brightening of local dark light features;

[0034] Step B22: Design an encoder, including two stages: s=0 and Stage s=1 , where Stage s=0 The Stage 2 consists of an illumination fusion attention mechanism, a two-dimensional selection scanning operation, and a multi-scale guided group attention module. s=1 It consists of an illumination fusion attention mechanism and a two-dimensional selective scanning operation; the illumination fusion attention mechanism used (J.Bai, Y.Yin, Q.He, Y.Li, X.Zhang, RetinexMamba: Retinex-based Mamba for low-light image enhancement, arXiv preprint arXiv, 2405.03349, 2024.) introduces illumination features for cross-attention operations, allowing the model to better focus on the dark areas that need to be enhanced; the two-dimensional selective scanning operation used (S.Wang, Q.Tao, and Z.Tang, Resvmunetx: A low-light enhancement network based on vmamba, arXiv preprint arXiv, 2401.10166, 2024.) includes a scan expansion operation, an S6 block, and a scan merging operation to solve the image dark light problem from a global dimension, wherein the S6 block represents that each feature in the sequence interacts with the previously scanned feature, and uses a compressed hidden state to reduce the square complexity to a linear complexity;

[0035] Step B23, design a feature reconstruction module based on center mask attention perception, which first performs pixel rearrangement operation, then performs center mask convolution operation, and finally performs mask attention operation, so as to learn how to reconstruct simulated noise features using limited non-noise information;

[0036] Step B24: Design an intermediate deep feature processing module, namely Stage s=2 , including an attention mechanism for illumination fusion, a two-dimensional selective scanning operation, and a feature reconstruction module for center mask attention perception.

[0037] Step B25: Design a decoder consisting of two stages: s=3 and Stage s=4 , where Stage s=3 It consists of an attention mechanism for illumination fusion and a two-dimensional selection scanning operation. s=4 It consists of an illumination fusion attention mechanism, a two-dimensional selective scanning operation, and a multi-scale guided grouping attention module in sequence.

[0038] In one embodiment of the present invention, the specific implementation steps of step B21 are as follows:

[0039] Step B211, design a multi-scale feature learning operation, including a three-branch depth-separable convolution operation, using different convolution kernel sizes in the three branches to achieve multi-scale local feature discovery. The specific formula is as follows:

[0040]

[0041] in, represents the features obtained after the attention mechanism of illumination fusion and two-dimensional selection scanning, s represents the sth stage, LP(·) represents the linear mapping operation, σ represents the GELU activation function, and DWConv c×c represents a depthwise separable convolution operation with a convolution kernel size of c×c, where c is 1, 3, or 5, and ∑ represents the summation of the output features of the three depthwise separable convolutions;

[0042] Step B212: Design a multi-branch multi-receptive field convolution operation to perform the features obtained in step B211 on the Perform four branch convolution operations;

[0043] First, a convolution block with a convolution kernel of e×f is composed of a two-dimensional convolution operation with a convolution kernel of e×f, a batch normalization operation, and a ReLU activation function. The convolution layer in the first branch is composed of a convolution block with a convolution kernel of 1×1. The convolution layers in the second branch, the third branch, and the fourth branch are composed of a convolution block with a convolution kernel of 1×1, a convolution block with a convolution kernel of 1×n, a convolution block with a convolution kernel of n×1, and an expanded convolution block with a convolution kernel of n×n. In the second branch, the third branch, and the fourth branch, n is 3, 5, and 7, respectively.

[0044] After the input feature P undergoes a multi-branch multi-receptive field convolution operation, the output features of the four branches are obtained. The output features obtained after splicing in the channel dimension are expressed as After that, a 2D convolution operation with a convolution kernel of 3×3 is used to reduce the channel dimension of X to make it consistent with the size of the input feature P, and finally a residual connection operation is performed with the input feature to obtain the output feature The specific process is as follows:

[0045] Y=Conv 3×3 (X)+Conv 1×1 (P)

[0046] Among them, Conv 3×3 (·) and Conv 1×1 Represent 3×3 convolution and 1×1 convolution operations respectively;

[0047] Step B213, design a grouping attention operation, including a grouping attention mechanism and a gated recurrent unit; first perform the grouping attention mechanism, which better distinguishes the bright and dark features from a global dimension by grouping the bright and dark features in a single image, and brightens the correctly grouped local dark features; first, randomly initialize the learnable clustering feature embedding Where S represents the number of groups in a single image, C represents the channel size, and then E is assigned to the initial clustering feature The clustering feature is derived from random initialization and is used to discover the grouping features of bright and dark light in the image, which is used to generate the first query feature, where t represents the number of updates of the learnable clustering feature; a total of three grouping attention mechanisms and gated recurrent units are required. After the tth attention mechanism and gated recurrent unit, the output feature is obtained by layer normalization and linear mapping operations to obtain a new query feature, which is passed to the next t+1th grouping attention mechanism and gated recurrent unit operation. After the last grouping attention mechanism and gated recurrent unit operation, the grouping information of bright and dark light in the image is propagated to the image features through the spatial position embedding operation to obtain the output

[0048] Specifically, first, position embedding is added to the multi-scale feature Y generated in step B212, and it is flattened into a 2D feature. Where N = H × W, then Y' is passed through a multi-layer perceptron and layer normalization to obtain feature A, while A is passed through layer normalization and linear mapping to generate a key matrix Sum Matrix At the same time t Generate query matrix through layer normalization and linear mapping The specific formula is as follows:

[0049] Q t =LN(y t )W q ,K=LN(A)W k,V=LN(A)W v

[0050] Among them, LN(·) represents layer normalization, W q , W k , W v Represents a linear mapping operation;

[0051] When the tth attention mechanism is performed, the query matrix Q is first t The dot product operation with the key matrix K is used to learn the relationship between different positions within the feature and obtain the attention weight Then, the weighted average operation is used to make the attention weight more stable; finally, in order to refer to the features of important positions, the weighted average attention weight is multiplied by the value matrix V to obtain the output features The specific operation process is as follows:

[0052]

[0053]

[0054] in, represents regularization, the superscript T of K represents the rank conversion operation, and Softmax(·) represents the normalized exponential function. Represents a weighted averaging operation;

[0055] Design a gated recurrent unit to update the learnable clustering feature representation; specifically, use the current clustering feature y t and feature O t , generating intermediate features through gated recurrent units The specific formula is as follows:

[0056] G t =GRU(O t ,y t )

[0057] Among them, GRU(·) represents the gated recurrent unit operation;

[0058] Finally, layer normalization and multi-layer perceptron are used to generate the t+1th updated clustering feature. The formula is as follows:

[0059] y t+1 =MLP(LN(G t ))+G t

[0060] Among them, LN(·) represents layer normalization and MLP represents multi-layer perceptron operation.

[0061] In one embodiment of the present invention, the specific implementation steps of step B23 are as follows:

[0062] Step B231, design a pixel rearrangement operation, by All pixels in the image are rearranged to break the correlation between local noises and make the noise distribution more random and uniform. After the pixel rearrangement operation, feature U is obtained. Specifically, feature U is first obtained through the reshaping operation. Then rearrange to obtain features Where r represents the reduction factor, and finally C×r 2 The size is The matrix is ​​arranged from left to right and from top to bottom and reshaped to obtain the output features

[0063] From the above rearrangement process, we can see that the local relationship of the noise is disrupted, making the local noise distribution after rearrangement random, making the subsequent feature reconstruction more robust when facing different noises;

[0064] Step B232, design a central mask convolution operation, use the limited features in the local range to perform feature reconstruction operation on the central mask, and increase the model's ability to reconstruct and remove local noise features; first, perform a mask operation on the central features in the local range; specifically, use the features at each pixel position Divide the center into a 3×3 matrix, perform mask operation on the center pixel, that is, assign a value of 0 and assign a value of 1 to the remaining pixels. The specific operation is as follows:

[0065]

[0066] Among them, variables m∈[i-1,i+1], n∈[j-1,j+1], when dividing the 3×3 matrix centered on the edge pixel feature, the pixel filling operation is performed on the vacant pixels on the edge, and * represents the dot multiplication operation. represents the center mask matrix centered at pixel i,j and of size 3×3;

[0067] After that, for H×W u" i,j The central mask matrix first performs a convolution operation with a convolution kernel of 3×3, and then performs two convolution operations with a convolution kernel of 1×1 in sequence to reconstruct the features of the central pixel of the mask and obtain the output feature B;

[0068] Step B233, design a window-based multi-head mask attention mechanism, first perform a random mask operation within the global feature range of the image; specifically, first randomly set the mask ratio p, which is randomly generated between 0.6 and 0.9, and then randomly select the pixels to be masked to ensure that the ratio of the number of pixels to be masked to the total number of pixels in the image is p, and randomly generate a sequence L of pixel positions to be masked in the image. If the pixel belongs to the sequence L, a mask operation is performed. If not, the pixel retains the original feature value. After the random mask operation, the output feature is obtained. The specific formula is as follows:

[0069]

[0070] Where (m,n) represents the pixel position with a height of m and a width of n in the image;

[0071] Then, the features are divided into h heads in the channel dimension for multi-head attention, and the query matrix is ​​obtained by linear mapping Key Matrix Sum Matrix Output features obtained through multi-head attention The specific formula is as follows:

[0072]

[0073] in, It is used for regularization operation, and R is used for position embedding operation;

[0074] Finally, after the reshaping operation, the size of feature A becomes And the multi-head features are spliced ​​to obtain

[0075] Step B234, design a masked attention operation to further enhance the model's ability to remove severe noise by increasing the range and amount of noise; Different from the center mask convolution operation in step B232, the masked attention operation in this step increases the amount and range of simulated noise, so that the model can better reconstruct and remove noise features when facing severe noise in the global range; First, the output feature of step B232 is Perform window-based segmentation operations to obtain segmented features Where Z represents the side length of the divided window; multiple window features are subjected to a window-based multi-head mask attention mechanism, and then the window merging operation is performed to obtain the intermediate features Finally, after layer normalization and multi-layer perceptron operation and residual connection, the output features are obtained. The specific formula is as follows:

[0076] Y=B+WM(MWA(LN(B')))

[0077] M=Y+MLP(LN(Y))

[0078] Among them, LN(·) represents the layer normalization operation, MWA(·) represents the window-based multi-head mask attention mechanism, MLP(·) represents the multi-layer perceptron operation, and WM(·) represents the window merging operation. Specifically, the shape is obtained after the first reshaping operation. The features of Z in the feature dimension are then reshaped by the second reshaping operation. 2 It is split into Z×Z and finally undergoes a third reshaping operation to obtain an output feature of shape H×W×C.

[0079] In one embodiment of the present invention, the specific implementation of step C is as follows:

[0080] Using L1 norm loss, update the network parameters. The specific formula is as follows:

[0081]

[0082] Among them, y i is the true value, The predicted value from the network output, N represents the number of samples.

[0083] In one embodiment of the present invention, step D is specifically implemented as follows:

[0084] The processed training data set obtained in step A is divided into J batches, where each batch contains Z pairs of images; for the zth low-light image I in the jth batch, the enhanced image I is obtained by the low-light image enhancement network based on illumination grouping and mask attention in step B. output ; By using the loss function designed in step C, the loss of the enhanced image is calculated to update the network parameters; the Adam optimizer is used to update the network parameters to obtain the trained low-light image model based on illumination grouping and mask attention.

[0085] The present invention also provides a computer-readable storage medium, on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the method steps described above can be implemented.

[0086] Compared with the prior art, the present invention has the following beneficial effects: First, in order to solve the problems of poor local detail information learning and local bright and dark light grouping in the Mamba method, the present invention designs a multi-scale guided grouping attention module, which includes a multi-scale feature learning operation, a multi-branch multi-receptive field convolution operation and a grouping attention mechanism. The multi-scale feature learning and the multi-branch multi-receptive field convolution operation can solve the problem of image detail loss from the perspective of local multi-receptive fields. The grouping attention mechanism improves the local brightness and restores the image contrast. In order to solve the noise problem in the brightening process, a center mask attention perception feature reconstruction module is designed, which includes a pixel rearrangement operation, a center mask convolution operation and a mask attention operation. The pixel rearrangement operation is used to interrupt the correlation between local noises and make the distribution of noise more random and uniform. The center mask convolution operation is used to train the model to learn to reconstruct the center mask using useful information in the local range, and the mask attention operation is used to enhance the model's noise removal ability when facing global multi-noise. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1 It is a flow chart for realizing the method of the present invention.

[0088] Figure 2 It is a structural diagram of a low-light image enhancement network based on illumination grouping and mask attention in an embodiment of the present invention.

[0089] Figure 3 is a structural diagram of an image brightening module in an embodiment of the present invention.

[0090] Figure 4 It is a structural diagram of a grouped attention module based on multi-scale guidance in an embodiment of the present invention.

[0091] Figure 5 It is a structural diagram of a center mask attention perception feature reconstruction module in an embodiment of the present invention. DETAILED DESCRIPTION

[0092] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.

[0093] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0094] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0095] The present invention provides a low-illumination image enhancement method based on illumination grouping and mask attention, comprising:

[0096] Step A: preprocessing the input image, including operations of image data pairing, data cropping and image enhancement, to obtain a training data set;

[0097] Step B, designing a low-light image enhancement network based on illumination grouping and mask attention, including an image brightening module, a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction, and a feature output module;

[0098] Step C, designing a loss function for updating network parameters of the low-light image enhancement network based on illumination grouping and mask attention in step B;

[0099] Step D: using the training data set to train the low-illumination image enhancement network based on illumination grouping and mask attention in step B to obtain a trained low-illumination image enhancement network based on illumination grouping and mask attention;

[0100] Step E: Use the trained low-illumination image enhancement network based on illumination grouping and mask attention obtained in step D to test the image and obtain a predicted normal illumination image.

[0101] The following is the specific implementation process of the present invention.

[0102] The present invention provides a low-light image enhancement method based on illumination grouping and mask attention, and the implementation flow chart is as follows: Figure 1 As shown in the network structure diagram Figure 2 As shown, the following steps are included:

[0103] Step A: preprocess the input image, including operations such as image data pairing, data cropping and image enhancement, to obtain a training data set;

[0104] Step B, designing a low-light image enhancement network based on illumination grouping and mask attention, the network comprising an image brightening module, a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction, and a feature output module;

[0105] Step C, designing a loss function for updating the network parameters in step B;

[0106] Step D: using the processed training data in step A to train the low-light image enhancement network in step B, to obtain a trained low-light image enhancement network based on illumination grouping and mask attention;

[0107] Step E: Use the trained low-illumination image enhancement network obtained in step D to test the image to obtain a predicted normal-illumination image.

[0108] Further, step A comprises the following steps:

[0109] Step A1, pairing the low-light image with the label image;

[0110] Step A2: randomly crop each low-light image and the label image to obtain an image of size H×W×3, where H and W are the height and width of the cropped image;

[0111] Step A3: Randomly use one of the following eight data augmentation methods on the paired low-light image and the label image: keep the original image, flip upside down, rotate 90 degrees counterclockwise, rotate 90 degrees counterclockwise and flip upside down, rotate 180 degrees counterclockwise, rotate 180 degrees counterclockwise and flip upside down, rotate 270 degrees counterclockwise, and rotate 270 degrees counterclockwise and flip upside down.

[0112] Further, step B comprises the following steps:

[0113] Step B1: Design an image brightening module, including two-dimensional convolution and depthwise separable convolution, to brighten low-light images. Generate preliminary brightened image features

[0114] Step B2: Design a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction. The overall architecture includes an encoder, a decoder, and an intermediate deep feature processing module. The module contains S stages, which are used to restore the initial brightened image feature obtained in step B1 to feature

[0115] Step B3: Design a feature output module, including a two-dimensional convolution operation and a residual connection operation, to generate a normal illumination image from the restored feature y' obtained in step B2

[0116] Furthermore, if Figure 3 As shown, step B1 includes the following steps:

[0117] Step B11: To obtain the brightening feature, the low-light image obtained in step A is Perform an averaging operation in the channel dimension to obtain the lighting prior features Where H and W represent the height and width of the image feature. input With the lighting prior feature I p After concatenation in the channel dimension, the size of the merged feature is H×W×4. Then, the brightening feature is obtained through two-dimensional convolution and depth-wise separable convolution operations. Where C represents the channel size of the feature, and the specific implementation method is:

[0118] Z lu =DWConv 5×5 (Conv 1×1 (Cat(I input ,I p )))

[0119] Among them, Cat(·) represents the merging operation, Conv 1×1 (·) indicates a two-dimensional convolution operation with a convolution kernel of 1×1, DWConv 5×5 (·) indicates a depthwise separable convolution operation with a kernel of 5×5;

[0120] Step B12: Brighten feature Z lu After two-dimensional convolution operation, and with the original image I input Perform multiplication and residual connection operations to generate a brightened image The specific formula is as follows:

[0121]

[0122] in, Represents element-wise multiplication, Conv 1×1 (·) indicates a two-dimensional convolution operation with a convolution kernel of 1×1;

[0123] Step B13: The brightened image I generated in step B12 is lu After two-dimensional convolution, we get feature T s , the specific formula is as follows:

[0124] T s =Conv 3×3 (I lu )

[0125] Among them, Conv 3×3 (·) indicates a two-dimensional convolution operation with a convolution kernel of 3×3;

[0126] Further, step B2 includes the following steps:

[0127] Step B21, design a multi-scale guided group attention module, which first undergoes a multi-scale feature learning operation, then a multi-branch multi-receptive field convolution operation, and finally a group attention operation, for local detail recovery and brightening of local dark light features;

[0128] Step B22: Design an encoder that includes two stages: s=0 and Stage s=1 , where Stage s=0 The Stage 2 consists of an illumination fusion attention mechanism, a two-dimensional selection scanning operation, and a multi-scale guided group attention module. s=1 It consists of an illumination fusion attention mechanism and a two-dimensional selective scanning operation. The illumination fusion attention mechanism used (J.Bai, Y.Yin, Q.He, Y.Li, X.Zhang, RetinexMamba: Retinex-based Mamba for low-light image enhancement, arXiv preprint arXiv, 2405.03349, 2024.) introduces illumination features for cross-attention operations, allowing the model to better focus on the dark areas that need to be enhanced. The two-dimensional selective scanning operation used (S.Wang, Q.Tao, and Z.Tang, Resvmunetx: A low-light enhancement network based on vmamba, arXiv preprint arXiv, 2401.10166, 2024.) contains a scan expansion operation, an S6 block, and a scan merging operation, which solves the image dark light problem from a global dimension. The S6 block represents that each feature in the sequence interacts with the previously scanned feature, and uses a compressed hidden state to reduce the square complexity to linear complexity;

[0129] Step B23, design a feature reconstruction module based on center mask attention perception. The module first undergoes pixel rearrangement operation, then center mask convolution operation, and finally mask attention operation, to learn how to reconstruct simulated noise features using limited non-noise information.

[0130] Step B24: Design an intermediate deep feature processing module, namely Stage s=2 , including the attention mechanism of illumination fusion, the two-dimensional selective scanning operation and the feature reconstruction module of center mask attention perception.

[0131] Step B25: Design a decoder consisting of two stages: s=3 and Stage s=4 , where Stages=3 It consists of an attention mechanism for illumination fusion and a two-dimensional selection scanning operation. s=4 It consists of an illumination fusion attention mechanism, a two-dimensional selective scanning operation, and a multi-scale guided grouping attention module in sequence.

[0132] Furthermore, if Figure 4 As shown, step B21 includes the following steps:

[0133] Step B211, design a multi-scale feature learning operation. This operation mainly includes a three-branch depth-separable convolution operation, and different convolution kernel sizes are used in the three branches to realize the discovery of multi-scale local features. The specific formula is as follows:

[0134]

[0135] in, represents the features obtained after the attention mechanism of illumination fusion and two-dimensional selection scanning, s represents the sth stage in the method. LP(·) represents the linear mapping operation, σ represents the GELU activation function, and DWConv c×c It represents a depth-wise separable convolution operation with a convolution kernel size of c×c. The value of c is 1, 3, or 5. '∑' represents the summation of the output features of the three depth-wise separable convolutions.

[0136] Step B212: Design a multi-branch multi-receptive field convolution operation. Four branch convolution operations are performed.

[0137] First, a convolution block with a convolution kernel of e×f is composed of a two-dimensional convolution operation with a convolution kernel of e×f, a batch normalization operation and a ReLU activation function. The convolution layer in the first branch is composed of a convolution block with a convolution kernel of 1×1, and the convolution layers in the second, third and fourth branches are composed of a convolution block with a convolution kernel of 1×1, a convolution block with a convolution kernel of 1×n, a convolution block with a convolution kernel of n×1 and an expanded convolution block with a convolution kernel of n×n. Among them, n in the second, third and fourth branches is 3, 5 and 7 respectively.

[0138] After the input feature P undergoes a multi-branch multi-receptive field convolution operation, the output features of the four branches are obtained. The output features obtained after splicing in the channel dimension are expressed as After that, a 2D convolution operation with a convolution kernel of 3×3 is used to reduce the channel dimension of X to make it consistent with the size of the input feature P, and finally a residual connection operation is performed with the input feature to obtain the output feature The specific process is as follows:

[0139] Y=Conv 3×3 (X)+Conv1×1 (P)

[0140] Among them, Conv 3×3 (·) and Conv 1×1 Represent 3×3 convolution and 1×1 convolution operations respectively.

[0141] Step B213, design a grouping attention operation. It mainly includes a grouping attention mechanism and a gated recurrent unit. First, the grouping attention mechanism is performed. The grouping attention mechanism performs a grouping operation on the bright and dark features in a single image to better distinguish the bright and dark features from a global dimension, and brighten the correctly grouped local dark features. First, randomly initialize the learnable clustering feature embedding Where S represents the number of groups in a single image, and C represents the channel size. Then assign E to the initial clustering feature This clustering feature is derived from random initialization and is used to discover the grouping features of bright and dark light in the image, which is used to generate the first query feature, where t represents the number of updates of the learnable clustering feature. This method requires three grouping attention mechanisms and gated recurrent units. After the tth attention mechanism and gated recurrent unit, the output feature is obtained by layer normalization and linear mapping operations to obtain a new query feature, which is passed to the next t+1th grouping attention mechanism and gated recurrent unit operation. After the last grouping attention mechanism and gated recurrent unit operation, the grouping information of bright and dark light in the image is propagated to the features in the image through the spatial position embedding operation to obtain the output

[0142] Specifically, first, position embedding is added to the multi-scale feature Y generated in step B212, and it is flattened into a 2D feature. Where N = H × W. Then Y' is passed through a multi-layer perceptron and layer normalization to obtain feature A, while A is passed through layer normalization and linear mapping to generate a key matrix Sum Matrix At the same time t Generate query matrix through layer normalization and linear mapping The specific formula is as follows:

[0143] Q t =LN(y t )W q ,K=LN(A)W k ,V=LN(A)W v

[0144] Among them, LN(·) represents layer normalization, W q , W k , W v Represents a linear map operation.

[0145] When the tth attention mechanism is performed, the query matrix Q is first t The dot product operation with the key matrix K is used to learn the relationship between different positions within the feature and obtain the attention weight Then, the weighted average operation is used to make the attention weight more stable. Finally, in order to refer to the features of important positions, the weighted average attention weight is multiplied by the value matrix V to obtain the output features. The specific operation process is as follows:

[0146]

[0147]

[0148] in, represents regularization, the superscript ‘T’ of K represents the rank-transfer operation, and Softmax(·) represents the normalized exponential function. Represents a weighted averaging operation.

[0149] Design a gated recurrent unit to update the learnable clustering feature representation. Specifically, use the current clustering feature y t and feature O t , generating intermediate features through gated recurrent units The specific formula is as follows:

[0150] G t =GRU(O t ,y t )

[0151] Here, GRU(·) represents the gated recurrent unit operation.

[0152] Finally, layer normalization and multi-layer perceptron are used to generate the t+1th updated clustering feature. The formula is as follows:

[0153] y t+1 =MLP(LN(G t ))+G t

[0154] Among them, LN(·) represents layer normalization and MLP represents multi-layer perceptron operation.

[0155] Furthermore, if Figure 5 As shown, step B23 includes the following steps:

[0156] Step B231, design a pixel rearrangement operation, by All pixels in the image are rearranged to break the correlation between local noises and make the distribution of noises more random and uniform. After the pixel rearrangement operation, feature U is obtained. Specifically, feature U is first obtained through the reshaping operation. Then rearrange to obtain features Where r represents the reduction factor, and finally 'C×r 2 'The size is The matrix is ​​arranged from left to right and from top to bottom and reshaped to obtain the output features

[0157] From the above rearrangement process, we can see that the local relationship of the noise is disrupted, making the local noise distribution after rearrangement random, making the subsequent feature reconstruction more robust when facing different noises;

[0158] Step B232, design the center mask convolution operation. Use the limited features in the local range to perform feature reconstruction on the center mask, and increase the model's ability to reconstruct and remove local noise features. First, perform a mask operation on the center feature in the local range. Specifically, use the feature at each pixel position as Divide the center into a 3×3 matrix, perform mask operation on the center pixel, that is, assign a value of '0' and assign a value of '1' to the remaining pixels. The specific operation is as follows:

[0159]

[0160] Among them, variables m∈[i-1,i+1], n∈[j-1,j+1], when dividing the 3×3 matrix centered on the edge pixel feature, the pixel filling operation is performed on the vacant pixels on the edge, and '*' represents the dot multiplication operation. Represents the center mask matrix centered at pixel i,j and of size 3×3.

[0161] After that, for the 'H×W' u" i,j The central mask matrix first performs a convolution operation with a convolution kernel of 3×3, and then performs two convolution operations with a convolution kernel of 1×1 in sequence to reconstruct the features of the central pixel of the mask and obtain the output feature B.

[0162] Step B233, design a window-based multi-head mask attention mechanism. First, perform a random mask operation within the global feature range of the image. Specifically, first randomly set the mask ratio p, which is randomly generated between 0.6 and 0.9, and then randomly select the pixels to be masked to ensure that the ratio of the number of pixels to be masked to the total number of pixels in the image is p. Randomly generate a sequence L of pixel positions that need to be masked in the image. If the pixel belongs to the sequence L, perform a mask operation. If not, the pixel retains the original feature value. After the random mask operation, the output feature is obtained. The specific formula is as follows:

[0163]

[0164] Where (m,n) represents the pixel position in the image with a height of m and a width of n.

[0165] Then, the features are divided into h heads in the channel dimension for multi-head attention, and the query matrix is ​​obtained by linear mapping Key Matrix Sum Matrix Output features obtained through multi-head attention The specific formula is as follows:

[0166]

[0167] in, It is used for regularization operation and R is used for position embedding operation.

[0168] Finally, after the reshaping operation, the size of feature A becomes And the multi-head features are spliced ​​to obtain

[0169] Step B234: Design a masked attention operation to further enhance the model's ability to remove severe noise by increasing the range and amount of noise. Unlike the center mask convolution operation in step B232, the masked attention operation in this step increases the amount and range of simulated noise, allowing the model to better reconstruct and remove noise features when facing severe noise in the global range. First, the output feature of step B232 is Perform window-based segmentation operations to obtain segmented features Where Z represents the side length of the partitioned window. Multiple window features are subjected to a window-based multi-head mask attention mechanism, and then the window merging operation is performed to obtain the intermediate features. Finally, after layer normalization and multi-layer perceptron operation and residual connection, the output features are obtained. The specific formula is as follows:

[0170] Y=B+WM(MWA(LN(B')))

[0171] M=Y+MLP(LN(Y))

[0172] Among them, LN(·) represents the layer normalization operation, MWA(·) represents the window-based multi-head mask attention mechanism, and MLP(·) represents the multi-layer perceptron operation. WM(·) represents the window merging operation. Specifically, the shape is obtained after the first reshaping operation. The features of , and then the second reshape operation to the feature dimension 'Z 2 ' is split into 'Z×Z', and finally the output features with the shape of 'H×W×C' are obtained after the third reshape operation.

[0173] Further, step C comprises the following steps:

[0174] Step C: Use L1 norm loss to update network parameters. The specific formula is as follows:

[0175]

[0176] Among them, y i is the true value, The predicted value from the network output, N represents the number of samples.

[0177] Further, step D comprises the following steps:

[0178] Step D: Training data set After the processed training set obtained in step A, it is divided into J batches, where each batch contains Z pairs of images; for the zth low-light image I of the jth batch, the enhanced image I is obtained by the low-light image enhancement network based on illumination grouping and mask attention in step B. output ; By using the loss function designed in step C, the loss of the enhanced image is calculated to update the network parameters; the Adam optimizer is used to update the network parameters to obtain the trained low-light image model based on illumination grouping and mask attention.

[0179] The present invention also provides a computer-readable storage medium, on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the method steps described above can be implemented.

[0180] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0181] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0182] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0183] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0184] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.

Claims

1. A low-light image enhancement method based on illumination grouping and mask attention, characterized in that: include: Step A: preprocessing the input image, including operations of image data pairing, data cropping and image enhancement, to obtain a training data set; Step B, designing a low-light image enhancement network based on illumination grouping and mask attention, including an image brightening module, a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction, and a feature output module; Step C, designing a loss function for updating network parameters of the low-light image enhancement network based on illumination grouping and mask attention in step B; Step D: using the training data set to train the low-illumination image enhancement network based on illumination grouping and mask attention in step B to obtain a trained low-illumination image enhancement network based on illumination grouping and mask attention; Step E: Use the trained low-illumination image enhancement network based on illumination grouping and mask attention obtained in step D to test the image and obtain a predicted normal illumination image.

2. The low-light image enhancement method based on illumination grouping and mask attention according to claim 1, characterized in that: The specific implementation steps of step A are as follows: Step A1, pairing the low-light image with the label image; Step A2: randomly crop each low-light image and the label image to obtain an image of size H×W×3, where H and W are the height and width of the cropped image; Step A3: Randomly use one of the following eight data augmentation methods on the paired low-light image and the label image: keep the original image, flip upside down, rotate 90 degrees counterclockwise, rotate 90 degrees counterclockwise and flip upside down, rotate 180 degrees counterclockwise, rotate 180 degrees counterclockwise and flip upside down, rotate 270 degrees counterclockwise, and rotate 270 degrees counterclockwise and flip upside down.

3. The low-light image enhancement method based on illumination grouping and mask attention according to claim 1, characterized in that: The specific implementation steps of step B are as follows: Step B1: Design an image brightening module, including two-dimensional convolution and depthwise separable convolution, to brighten low-light images. Generate preliminary brightened image features Where H and W represent the height and width of the image feature, and C represents the channel of the image feature; Step B2: Design a degradation recovery module based on multi-scale illumination grouping and mask attention feature reconstruction. The overall architecture includes an encoder, a decoder, and an intermediate deep feature processing module. The degradation recovery module includes S stages, which are used to restore the initial brightened image features obtained in step B1 to feature Step B3: Design a feature output module, including a two-dimensional convolution operation and a residual connection operation, to generate a normal illumination image from the restored feature y' obtained in step B2 4. The low-light image enhancement method based on illumination grouping and mask attention according to claim 3, characterized in that: The specific implementation steps of step B1 are as follows: Step B11: To obtain brightening features, the low-light image in the training data set obtained in step A is Perform an averaging operation in the channel dimension to obtain the lighting prior features Then the low-light image I input With the lighting prior feature I p After concatenation in the channel dimension, the size of the merged feature is H×W×4, and then the brightening feature is obtained through two-dimensional convolution operation and depth-separable convolution operation. The specific implementation method is: Z lu =DWConv 5×5 (Conv 1×1 (Cat(I input ,I p ))) Among them, Cat(·) represents the merging operation, Conv 1×1 (·) indicates a two-dimensional convolution operation with a convolution kernel of 1×1, DWConv 5×5 (·) indicates a depthwise separable convolution operation with a kernel of 5×5; Step B12: Brighten feature Z lu After two-dimensional convolution operation, and with the original image I input Perform multiplication and residual connection operations to generate a brightened image The specific formula is as follows: in, Represents element-wise multiplication, Conv 1×1 (·) indicates a two-dimensional convolution operation with a convolution kernel of 1×1; Step B13: The brightened image I generated in step B12 is lu After two-dimensional convolution, we get feature T s , the specific formula is as follows: T s =Conv 3×3 (I lu ) Among them, Conv 3×3 (·) represents a two-dimensional convolution operation with a convolution kernel of 3×3.

5. The low-light image enhancement method based on illumination grouping and mask attention according to claim 3, characterized in that: The specific implementation steps of step B2 are as follows: Step B21, design a multi-scale guided group attention module, which first undergoes a multi-scale feature learning operation, then a multi-branch multi-receptive field convolution operation, and finally a group attention operation, for local detail recovery and brightening of local dark light features; Step B22: Design an encoder, including two stages: s=0 and Stage s=1 , where Stage s=0 The Stage 2 consists of an illumination fusion attention mechanism, a two-dimensional selection scanning operation, and a multi-scale guided group attention module. s=1 It consists of an attention mechanism for illumination fusion and a 2D selective scanning operation; The illumination fusion attention mechanism used introduces illumination features for cross-attention operations, allowing the model to better focus on dark areas that need to be enhanced. The two-dimensional selective scanning operation used includes scan expansion operation, S6 block and scan merging operation, which solves the image dark problem from a global dimension. The S6 block represents the interaction between each feature in the sequence and the previously scanned feature, and uses a compressed hidden state to reduce the square complexity to linear complexity. Step B23, design a feature reconstruction module based on center mask attention perception, which first performs pixel rearrangement operation, then performs center mask convolution operation, and finally performs mask attention operation, so as to learn how to reconstruct simulated noise features using limited non-noise information; Step B24: Design an intermediate deep feature processing module, namely Stage s=2 , including an attention mechanism for illumination fusion, a two-dimensional selective scanning operation, and a feature reconstruction module for center mask attention perception. Step B25: Design a decoder consisting of two stages: s=3 and Stage s=4 , where Stage s=3 It consists of an attention mechanism for illumination fusion and a two-dimensional selection scanning operation. s=4 It consists of an illumination fusion attention mechanism, a two-dimensional selective scanning operation, and a multi-scale guided grouping attention module in sequence.

6. The low-light image enhancement method based on illumination grouping and mask attention according to claim 5, characterized in that: The specific implementation steps of step B21 are as follows: Step B211, design a multi-scale feature learning operation, including a three-branch depth-separable convolution operation, using different convolution kernel sizes in the three branches to achieve multi-scale local feature discovery. The specific formula is as follows: in, represents the features obtained after the attention mechanism of illumination fusion and two-dimensional selection scanning, s represents the sth stage, LP(·) represents the linear mapping operation, σ represents the GELU activation function, and DWConv c×c represents a depthwise separable convolution operation with a convolution kernel size of c×c, where c is 1, 3, or 5, and ∑ represents the summation of the output features of the three depthwise separable convolutions; Step B212: Design a multi-branch multi-receptive field convolution operation to perform the features obtained in step B211 on the Perform four branch convolution operations; First, a convolution block with a convolution kernel of e×f is composed of a two-dimensional convolution operation with a convolution kernel of e×f, a batch normalization operation, and a ReLU activation function. The convolution layer in the first branch is composed of a convolution block with a convolution kernel of 1×1. The convolution layers in the second branch, the third branch, and the fourth branch are composed of a convolution block with a convolution kernel of 1×1, a convolution block with a convolution kernel of 1×n, a convolution block with a convolution kernel of n×1, and an expanded convolution block with a convolution kernel of n×n. In the second branch, the third branch, and the fourth branch, n is 3, 5, and 7, respectively. After the input feature P undergoes a multi-branch multi-receptive field convolution operation, the output features of the four branches are obtained. The output features obtained after splicing in the channel dimension are expressed as After that, a 2D convolution operation with a convolution kernel of 3×3 is used to reduce the channel dimension of X to make it consistent with the size of the input feature P, and finally a residual connection operation is performed with the input feature to obtain the output feature The specific process is as follows: Y=Conv 3×3 (X)+Conv 1×1 (P) Among them, Conv 3×3 (·) and Conv 1×1 Represent 3×3 convolution and 1×1 convolution operations respectively; Step B213, design a grouping attention operation, including a grouping attention mechanism and a gated recurrent unit; first perform the grouping attention mechanism, which better distinguishes the bright and dark features from a global dimension by grouping the bright and dark features in a single image, and brightens the correctly grouped local dark features; first, randomly initialize the learnable clustering feature embedding Where S represents the number of groups in a single image, C represents the channel size, and then E is assigned to the initial clustering feature The clustering feature is derived from random initialization and is used to discover the grouping features of bright and dark light in the image, which is used to generate the first query feature, where t represents the number of updates of the learnable clustering feature; a total of three grouping attention mechanisms and gated recurrent units are required. After the tth attention mechanism and gated recurrent unit, the output feature is obtained by layer normalization and linear mapping operations to obtain a new query feature, which is passed to the next t+1th grouping attention mechanism and gated recurrent unit operation. After the last grouping attention mechanism and gated recurrent unit operation, the grouping information of bright and dark light in the image is propagated to the image features through the spatial position embedding operation to obtain the output Specifically, first, position embedding is added to the multi-scale feature Y generated in step B212, and it is flattened into a 2D feature. Where N = H × W, then Y' is passed through a multi-layer perceptron and layer normalization to obtain feature A, while A is passed through layer normalization and linear mapping to generate a key matrix Sum Matrix At the same time t Generate query matrix through layer normalization and linear mapping The specific formula is as follows: Q t =LN(y t )W q ,K=LN(A)W k ,V=LN(A)W v Among them, LN(·) represents layer normalization, W q , W k , W v Represents a linear mapping operation; When the tth attention mechanism is performed, the query matrix Q is first t The dot product operation with the key matrix K is used to learn the relationship between different positions within the feature and obtain the attention weight Then, the weighted average operation is used to make the attention weight more stable; finally, in order to refer to the features of important positions, the weighted average attention weight is multiplied by the value matrix V to obtain the output features The specific operation process is as follows: in, represents regularization, the superscript T of K represents the rank conversion operation, and Softmax(·) represents the normalized exponential function. Represents a weighted averaging operation; Design a gated recurrent unit to update the learnable clustering feature representation; specifically, use the current clustering feature y i and feature O t , generating intermediate features through gated recurrent units The specific formula is as follows: G t =GRU(O t ,y t ) Among them, GRU(·) represents the gated recurrent unit operation; Finally, layer normalization and multi-layer perceptron are used to generate the t+1th updated clustering feature. The formula is as follows: y t+1 =MLP(LN(G t ))+G t Among them, LN(.) represents layer normalization and MLP represents multi-layer perceptron operation.

7. The low-light image enhancement method based on illumination grouping and mask attention according to claim 5, characterized in that: The specific implementation steps of step B23 are as follows: Step B231, design a pixel rearrangement operation, by All pixels in the image are rearranged to break the correlation between local noises and make the noise distribution more random and uniform. After the pixel rearrangement operation, feature U is obtained. Specifically, feature U is first obtained through the reshaping operation. Then rearrange to obtain features Where r represents the reduction factor, and finally C×r 2 The size is The matrix is ​​arranged from left to right and from top to bottom and reshaped to obtain the output features Step B232, design a central mask convolution operation, use the limited features in the local range to perform feature reconstruction operation on the central mask, and increase the model's ability to reconstruct and remove local noise features; first, perform a mask operation on the central features in the local range; specifically, use the features at each pixel position Divide the center into a 3×3 matrix, perform mask operation on the center pixel, that is, assign a value of 0 and assign a value of 1 to the remaining pixels. The specific operation is as follows: Among them, variables m∈[i-1,i+1], n∈[j-1,j+1], when dividing the 3×3 matrix centered on the edge pixel feature, the pixel filling operation is performed on the vacant pixels on the edge, and * represents the dot multiplication operation. represents the center mask matrix centered at pixel i,j and of size 3×3; After that, for H×W u" i,j The central mask matrix first performs a convolution operation with a convolution kernel of 3×3, and then performs two convolution operations with a convolution kernel of 1×1 in sequence to reconstruct the features of the central pixel of the mask and obtain the output feature B; Step B233, design a window-based multi-head mask attention mechanism, first perform a random mask operation within the global feature range of the image; specifically, first randomly set the mask ratio p, which is randomly generated between 0.6 and 0.9, and then randomly select the pixels to be masked to ensure that the ratio of the number of pixels to be masked to the total number of pixels in the image is p, and randomly generate a sequence L of pixel positions to be masked in the image. If the pixel belongs to the sequence L, a mask operation is performed. If not, the pixel retains the original feature value. After the random mask operation, the output feature is obtained. The specific formula is as follows: Where (m,n) represents the pixel position with a height of m and a width of n in the image; Then, the features are divided into h heads in the channel dimension for multi-head attention, and the query matrix is ​​obtained by linear mapping Key Matrix Sum Matrix Output features obtained through multi-head attention The specific formula is as follows: in, It is used for regularization operation, and R is used for position embedding operation; Finally, after the reshaping operation, the size of feature A becomes And the multi-head features are spliced ​​to obtain Step B234, design a masked attention operation to further enhance the model's ability to remove severe noise by increasing the range and amount of noise; Different from the center mask convolution operation in step B232, the masked attention operation in this step increases the amount and range of simulated noise, so that the model can better reconstruct and remove noise features when facing severe noise in the global range; First, the output feature of step B232 is Perform window-based segmentation operations to obtain segmented features Where Z represents the side length of the divided window; multiple window features are subjected to a window-based multi-head mask attention mechanism, and then the window merging operation is performed to obtain the intermediate features Finally, after layer normalization and multi-layer perceptron operation and residual connection, the output features are obtained. The specific formula is as follows: Y=B+WM(MWA(LN(B'))) M=Y+MLP(LN(Y)) Among them, LN(·) represents the layer normalization operation, MWA(·) represents the window-based multi-head mask attention mechanism, MLP(·) represents the multi-layer perceptron operation, and WM(·) represents the window merging operation. Specifically, the shape is obtained after the first reshaping operation. The features of Z in the feature dimension are then reshaped by the second reshaping operation. 2 It is split into Z×Z and finally undergoes a third reshaping operation to obtain an output feature of shape H×W×C.

8. The low-light image enhancement method based on illumination grouping and mask attention according to claim 1, characterized in that: The specific implementation of step C is: Using L1 norm loss, update the network parameters. The specific formula is as follows: Among them, y i is the true value, The predicted value from the network output, N represents the number of samples.

9. The low-light image enhancement method based on illumination grouping and mask attention according to claim 1, characterized in that: The step D is specifically implemented as follows: The processed training data set obtained in step A is divided into J batches, where each batch contains Z pairs of images; for the zth low-light image I in the jth batch, the enhanced image I is obtained by the low-light image enhancement network based on illumination grouping and mask attention in step B. output ; By using the loss function designed in step C, the loss of the enhanced image is calculated to update the network parameters; The Adam optimizer is used to update the network parameters to obtain the trained low-light image model based on illumination grouping and mask attention.

10. A computer-readable storage medium having stored thereon computer program instructions that can be executed by a processor, and when the processor executes the computer program instructions, the method steps according to any one of claims 1 to 9 can be implemented.

Citation Information

Patent Citations

  • Low-illumination image enhancement method based on Retinex and deep learning

    CN111968044A

  • Low-illumination image enhancement method, system and device and medium

    CN114359073A

  • Low-illumination image enhancement method based on recursive interactive attention

    CN116309182A

  • Low-illumination image restoration network containing large number of zero element pixels and restoration method

    CN119151833A

  • Low-light image enhancement method based on deep Retinex

    JP7493867B1