A low-light image enhancement method based on recursive interactive attention
By introducing a recursive interactive attention mechanism in the low-illumination image enhancement method, combining local and global attention, the problem of poor image enhancement effect in the prior art is solved, and a more natural and detailed image enhancement effect is achieved.
Patent Information
- Application Number
- CN202310356275.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-06
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-04-06
AI Technical Summary
The existing low-illumination image enhancement methods are difficult to effectively process images with uneven light distribution, resulting in insufficient sense of hierarchy of the enhancement result, and easy introduction of artifacts and noise, affecting image quality.
The low-illumination image enhancement method based on recursive interactive attention is adopted, combining the dual attention of local block perception and the offset attention of global context perception, and the expressive ability and stability of the model are enhanced through recursive and interactive strategies, fully considering the local details and global structure of the image.
It achieves a more natural and detailed image enhancement effect, improves the quality and brightness of the image, avoids problems such as excessive enhancement, artifacts and details, and adapts to different low-illumination image scenes.
Smart Images

Figure CN116309182B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the fields of image and video processing and computer vision technology, and in particular to a low-illumination image enhancement method based on recursive interactive attention. Background Art
[0002] With the continuous advancement of science and technology, the replacement and popularization of shooting and display devices such as smart phones and digital cameras, more and more image data in life exist in digital form. These data include photos taken in daily life, medical images and other image data from all walks of life. Among these data, due to the limitations of shooting environment and equipment, the proportion of low-light images continues to increase. Low-light images are dim in brightness, low contrast, color distortion, blurred image details and other problems due to insufficient light when shooting. These problems greatly affect the visual effect of the image. Low-light image enhancement aims to improve the quality of low-light images through digital image processing technology, making low-light images clearer, brighter, more colorful and rich in details. This can not only improve the visual effect of the image, but also improve the visualization effect of the image, bringing more convenience and comfort to people.
[0003] Low-light image enhancement has a wide range of applications in various fields, including security monitoring, transportation, night photography, etc. In the field of security monitoring, due to insufficient light in the night environment, traditional surveillance cameras cannot obtain clear images, which affects the monitoring effect. By enhancing the low-light image, the clarity and brightness of night monitoring can be improved, and the detailed information of the image can be enhanced, so as to better play the role of monitoring. In the field of transportation, when identifying a moving vehicle, low-light image enhancement can make the outline of the vehicle clearer and the edge more obvious, thereby improving the accuracy of vehicle identification. In addition, when identifying pedestrians, low-light image enhancement can also enhance the outline of pedestrians and reduce background noise, thereby improving the accuracy of pedestrian identification. In the field of night photography, due to insufficient light in the night environment, the photos taken are often dim and the color is distorted. Low-light image enhancement technology can enhance the brightness and color of the image, making the photos taken clearer and more realistic.
[0004] Early low-light image enhancement methods were mainly based on histogram equalization and Retinex theory. The method based on histogram equalization assumes that under normal illumination, the pixel values of the image are evenly distributed in all possible gray levels, so that the image will show high contrast and wide dynamic range. The global histogram equalization method has low computational complexity and high processing efficiency, and is suitable for low-light images with relatively uniform overall illumination distribution. However, this type of method counts the grayscale value of the entire image and lacks consideration of the intensity relationship between adjacent pixels. Therefore, for images with uneven illumination distribution, it is difficult to restore some local areas to the optimal value, resulting in insufficient layering of the enhancement results. Although the local histogram equalization method proposed to address this problem can effectively enhance the brightness and contrast of low-light images, it still does not specifically deal with the potential serious noise interference in low-light images, and may even amplify the noise. In addition, enhancing the contrast through simple function mapping is prone to color distortion. The method based on Retinex theory represents the low-light image as the product of the illumination component and the reflection component. The illumination component depends on the characteristics of the ambient light and determines the dynamic range of the image, while the reflection component is independent of the illumination and reflects the inherent properties of the object. Due to its physically explainable theoretical basis, methods based on Retinex theory can usually achieve good enhancement effects, but are limited by the number of parameters and the complexity of the function, and have limitations in the decomposition of the reflection component and the illumination component, which can easily lead to overexposure or underexposure of the image enhancement results, which are quite different from the real image.
[0005] With the development of deep learning, benefiting from the powerful feature learning ability of convolutional neural networks, low-light image enhancement methods based on deep learning have been widely studied and developed in recent years. For example, the combination of Retinex theory and deep learning methods can significantly improve image brightness, and deep learning methods that simulate image conversion can better enhance color and brightness. However, these methods usually have the following problems: first, it is easy to cause excessive enhancement of the image, affecting the image quality; second, artifacts are often introduced in the enhanced image, affecting the visual effect; third, there is a lack of targeted processing of image details, resulting in serious noise and detail loss in the enhanced image. Summary of the invention
[0006] In view of this, the purpose of the present invention is to provide a low-light image enhancement method based on recursive interactive attention. The method combines the local block-aware dual attention and the global context-aware shiftable attention, which can fully consider the local details and global structure of the image, making the enhanced image more natural and detailed; the introduced recursive and interactive strategies can also enhance the expressiveness and stability of the model, so that the model can better adapt to different low-light image scenes.
[0007] To achieve the above object, the present invention adopts the following technical solution: a low-light image enhancement method based on recursive interactive attention, comprising the following steps:
[0008] Step A: preprocess the input image, including image pairing, cropping, and data enhancement, to obtain a training data set;
[0009] Step B: designing a recursive interactive attention enhancement network, which consists of an input mapping module, a recursive interactive attention enhancement network and an output mapping module;
[0010] Step C, designing a loss function for training the network designed in step B;
[0011] Step D: training the recursive interactive attention enhancement network using the training dataset;
[0012] Step E: input the image to be tested into the network, and use the trained network to generate a normal illumination image.
[0013] In a preferred embodiment, the specific implementation steps of step A are as follows:
[0014] Step A1: pairing the normal illumination image with the low illumination image as a label image;
[0015] Step A2: randomly crop each low-light image of size H×W×3 into an image of size P×P×3, and use the same random cropping method for its corresponding normal-light image to ensure that they have the same size and position, where H and W are the height and width of the low-light image and the normal-light image, and P is the height and width of the cropped image;
[0016] Step A3: For each training paired image, randomly apply one of the following eight data augmentation methods: keep the original image, flip vertically, rotate 90 degrees, rotate 90 degrees and then flip vertically, rotate 180 degrees, rotate 180 degrees and then flip vertically, rotate 270 degrees, and rotate 270 degrees and then flip vertically.
[0017] In a preferred embodiment, the specific implementation steps of step B are as follows:
[0018] Step B1, designing an input mapping module, including a convolution layer, an activation function and a block coding layer, for extracting features from the input low-light image, dividing it into blocks and converting it into sequence-level features;
[0019] Step B2: Design a recursive interactive attention enhancement network, which consists of L recursive interaction blocks whose parameters are not shared, to achieve deep enhancement of features;
[0020] Step B3, design an output mapping module, including a block decoding layer and a convolutional layer, to reorganize the output features of the recursive interactive attention enhancement network into complete image features, and project the features into an image space with 3 channels;
[0021] The specific implementation steps of step B1 are as follows:
[0022] Step B11: A convolutional layer is formed by a convolution kernel of size 3×3, and LeakyReLU is used as the activation function to transform the input image Perform feature extraction, where H and W are the height and width of the low-light image respectively;
[0023] Step B12, a convolution layer with a convolution kernel size of D×D and a step size of D constitutes the core operation of the block coding layer. This operation divides the features obtained in step B11 into N blocks of size D×D, that is, The channel dimension of each block is C; then the features of these blocks are flattened, and the resulting sequence-level features As the input of the first recursive interaction block; the specific formula is as follows:
[0024] X (0) =PatchEmbed(LeakyReLU(Conv3(I in )))
[0025] Among them, Conv3 represents a 3×3 convolutional layer, LeakyReLU(·) represents the activation function, and PatchEmbed(·) represents the block coding layer operation.
[0026] In a preferred embodiment, the specific implementation steps of step B2 are as follows:
[0027] Step B21, designing recursive interaction blocks, each recursive interaction block includes T parameter-sharing recursive interaction units;
[0028] Step B22: The feature X obtained in step B12 (0) As the input of the first recursive interaction block, after l recursive interaction blocks and T recursions, the output feature is obtained Where l∈{1,2,...,L}; then, the output of the last recursive interaction block after T recursions is expressed as The specific formula is as follows:
[0029]
[0030] Among them, RIB l represents the lth recursive interaction block, Indicates the stacking process of the network.
[0031] In a preferred embodiment, the specific implementation steps of step B21 are as follows:
[0032] Step B211, designing a recursive interaction unit, which consists of a local block-aware dual attention module, a global context-aware shiftable attention module, an interaction operation between the two modules, and two operations: sequence-level feature reshaping and block-level feature reshaping;
[0033] Step B212: As the input feature of the lth recursive interaction block, the output feature of the current recursive interaction block is obtained after passing through T parameter-sharing recursive interaction units. The specific formula is as follows:
[0034]
[0035] Among them, RIU t represents the t-th recursive interaction unit, t∈{1,2,...,T}, Indicates repeated processing of the network, that is, parameter reuse.
[0036] In a preferred embodiment, the specific implementation steps of step B211 are as follows:
[0037] Step B2111, design a local block-aware dual attention module, the core of which includes a convolution unit, a channel attention unit and a spatial attention unit; for the input features First reshape it into block-level features Then, feature embedding is performed through a convolution unit, which is composed of a 3×3 convolution layer, a ReLu activation function, and a 3×3 convolution layer stacked in sequence. The mapped feature is represented as The specific formula is as follows:
[0038]
[0039]
[0040] Among them, Reshape1(·) is a reshaping operation that converts the N×C sequence-level features into The block-level features of , ConvUnit(·) represents the convolution unit;
[0041] Step B2112: Change the features of step B2211 The channel attention unit and the spatial attention unit are sent to obtain the weighted features of the channel attention and spatial attention respectively, and then the two features are concatenated, and a convolution layer composed of a convolution kernel of size 1×1 is used to map the features to a low-dimensional space, and finally Perform residual connection operation to obtain the output features of the local block-aware dual attention module The specific formula is as follows:
[0042]
[0043] Among them, CA represents the channel attention unit, SA represents the spatial attention unit, [·;·] represents the feature concatenation operation, and Conv1 represents the 1×1 convolutional layer;
[0044] Step B2113: Design the interactive operation from the local block-aware dual attention module to the global context-aware shiftable attention module. For the features output by the local block-aware dual attention module Reshape it back to sequence-level features and then compare it with the features Fusion as input to the global context-aware shiftable attention module The specific formula is:
[0045]
[0046] Reshape1(·) is a reshaping operation. The conversion of N×C sequence-level features into block-level features;
[0047] Step B2114, design a global context-aware shiftable attention module, the core of which includes layer normalization operation, shiftable multi-head self-attention and multi-layer perceptron; for the input features The features are then sent to the normalization layer and the offset multi-head self-attention layer, and the obtained features are then combined with Residual connection, output intermediate features Then, The features are sent to the normalization layer and the multi-head perceptron in turn, and then Residual connection, outputting the final features of the global context-aware shiftable attention module The specific formula is as follows:
[0048]
[0049]
[0050] Among them, LN(·) represents the layer normalization operation, MHSA represents the shiftable multi-head self-attention, and MLP represents the multi-head perceptron;
[0051] Step B2115: Design the interactive operation of the global context-aware shiftable attention module to the local block-aware dual attention module. For the features output by the global context-aware shiftable attention module Reshape it back to block-level features and then compare it with the features The output of the dual attention module is fused to update the local patch-aware The specific formula is:
[0052]
[0053] Among them, Reshape1(·) means reshaping the sequence-level features into block-level features.
[0054] In a preferred embodiment, the specific implementation of step B2112 is as follows:
[0055] Step B21121: Design the channel attention unit. First, input the feature The features are compressed into a global average pooling layer to obtain channel descriptors. Next, two 1×1 convolutional layers are used with a ReLu activation function in between to reduce and increase the dimensionality. Finally, the channel weights are limited between 0 and 1 through the Sigmoid activation function to obtain the channel attention weights. Finally, through the multiplication operation With channel weight α c Multiply them together to get the channel attention weighted features The specific formula is:
[0056]
[0057]
[0058] Among them, Avgpool(·) represents the global average pooling operation, Conv1 represents the 1×1 convolution layer, ReLu(·) is the activation function, σ(·) is the Sigmoid function, represents element-wise multiplication;
[0059] Step B21122: Design the spatial attention unit. First, input features Perform global maximum pooling and global average pooling operations along the channel dimension to obtain the maximum pooling feature map and the average pooling feature map, and then concatenate them along the channel dimension; then use a 3×3 convolution layer to map the feature map back to the 1-dimensional feature of the channel, and then use the Sigmoid activation function to limit the weight between 0 and 1 to obtain the spatial attention weight Finally, through the multiplication operation With the spatial attention weight α s Multiply them together to get the spatial attention weighted features The specific formula is:
[0060]
[0061]
[0062] Among them, AvgPool(·) represents the global average pooling operation, MaxPool(·) represents the global maximum pooling operation, [·;·] represents the concatenation operation, Conv3 represents the 3×3 convolutional layer, σ(·) is the Sigmoid function, Represents element-wise multiplication.
[0063] In a preferred embodiment, the specific implementation steps of step B3 are as follows:
[0064] Step B31, the core operation of the block decoding layer is composed of a bilinear upsampling layer and a 3×3 convolutional layer. This operation converts the N×C size feature of the last recursive interaction block output obtained in step B22 into Reshape to The block-level features of size are then upsampled to H×W×C features by a bilinear upsampling layer, and finally mapped by a 3×3 convolutional layer to obtain the features with reduced dimension.
[0065] Step B32: For the features output from step B31, a 3×3 convolutional layer is used to map them back to the image space and then combined with the input features I in The residual connection obtains the enhanced image I en ; The specific formula is as follows:
[0066]
[0067] Among them, PatchUnEmbed(·) is a block decoding operation, and Conv3 represents a 3×3 convolutional layer.
[0068] In a preferred embodiment, the specific implementation of step C is as follows:
[0069] Step C: Design the loss function, using Smooth L1 loss and VGG perceptual loss l perceptual Composition, the total objective loss function of the network l total It is expressed as follows:
[0070]
[0071]
[0072] lperceptual =||(VGG 3,8,15 (I en )-VGG 3,8,15 (I gt ))|| 2
[0073] Where λ is a balance parameter, E=I en -I gt , I gt is an image with normal illumination; ||·|| 2 Indicates the calculation of mean square error, VGG 3,8,15 (·) indicates that the features of layer 3, layer 8, and layer 15 are extracted using the VGG-16 classification model pre-trained on the ImageNet dataset.
[0074] In a preferred embodiment, the specific implementation steps of step D are as follows:
[0075] Step D1, randomly divide the training data set obtained in step A into several batches, each batch containing N pairs of images;
[0076] Step D2: input low-light image I in , after the recursive interactive attention enhancement network in step B, the enhanced image I is obtained en , use the formula in step C to calculate the loss l total ;
[0077] Step D3: Calculate the gradient of the parameters in the network using the back propagation method according to the loss, and update the network parameters using the Adam optimization method;
[0078] Step D4: Repeat steps D1 to D3 in batches to obtain a recursive interactive attention enhancement network model.
[0079] Compared with the prior art, the present invention has the following beneficial effects: First, the present invention encodes features in the input mapping module, improves the feature expression ability of the image, and can better retain the detail information of the image. Secondly, the present invention designs a recursive interactive attention enhancement network, including a local block-aware dual attention module and a global context-aware shiftable attention module. These modules regulate the outflow of effective information through spatial and channel attention mechanisms, can more accurately suppress problems such as artifacts and over-enhancement, and obtain better image enhancement effects. In addition, the present invention designs recursive and interactive operations to promote global consistency representation, can better improve the quality and brightness of the image, and avoid problems such as over-enhancement, artifacts and detail loss. Finally, the present invention uses an output mapping module to convert the enhanced feature map into the final image, which can better retain the detail information of the image and improve the quality and brightness of the image. Different from other recent low-light image enhancement methods based on Transformer, the present invention not only combines the advantages of local block perception and global context perception, but also uses only a small number of recursive interactive blocks, avoids the use of complex network structures such as large Transformers, and can achieve efficient low-light image enhancement under limited computing resources and storage space, reducing the implementation cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 It is a flow chart of the implementation of the method in the preferred embodiment of the present invention.
[0081] Figure 2 It is a structural diagram of a recursive interactive enhancement network in a preferred embodiment of the present invention.
[0082] Figure 3 It is a recursive interaction block structure diagram in a preferred embodiment of the present invention.
[0083] Figure 4 It is a structural diagram of the local block-aware dual attention module in a preferred embodiment of the present invention.
[0084] Figure 5 It is a structural diagram of the global context-aware shiftable attention module in the preferred embodiment of the present invention. DETAILED DESCRIPTION
[0085] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0086] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0087] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.
[0088] The present invention provides a low-light image enhancement method based on recursive interactive attention, such as Figure 1-5 As shown, the following steps are included:
[0089] Step A: preprocess the input image, including image pairing, cropping, and data enhancement, to obtain a training data set;
[0090] Step B, designing a recursive interactive attention enhancement network, which consists of an input mapping module, a recursive interactive attention enhancement network and an output mapping module;
[0091] Step C, designing a loss function for training the network designed in step B;
[0092] Step D: training the recursive interactive attention enhancement network using the training dataset;
[0093] Step E: input the image to be tested into the network, and use the trained network to generate a normal illumination image.
[0094] Furthermore, the step A comprises the following steps:
[0095] Step A1: pairing the normal illumination image with the low illumination image as a label image;
[0096] Step A2: randomly crop each low-light image of size H×W×3 into an image of size P×P×3, and use the same random cropping method for its corresponding normal-light image to ensure that they have the same size and position, where H and W are the height and width of the low-light image and the normal-light image, and P is the height and width of the cropped image;
[0097] Step A3: For each training paired image, randomly apply one of the following eight data augmentation methods: keep the original image, flip vertically, rotate 90 degrees, rotate 90 degrees and then flip vertically, rotate 180 degrees, rotate 180 degrees and then flip vertically, rotate 270 degrees, and rotate 270 degrees and then flip vertically.
[0098] Furthermore, the step B comprises the following steps:
[0099] Step B1, designing an input mapping module, including a convolution layer, an activation function and a block coding layer, for extracting features from the input low-light image, dividing it into blocks and converting it into sequence-level features;
[0100] Step B2: Design a recursive interactive attention enhancement network, which consists of L recursive interaction blocks whose parameters are not shared, to achieve deep enhancement of features;
[0101] Step B3: Design an output mapping module, including a block decoding layer and a convolutional layer, to reorganize the output features of the recursive interactive attention enhancement network into complete image features and project the features into an image space with 3 channels.
[0102] Furthermore, the step B1 comprises the following steps:
[0103] Step B11: A convolutional layer is formed by a convolution kernel of size 3×3, and LeakyReLU is used as the activation function to transform the input image Perform feature extraction, where H and W are the height and width of the low-light image respectively;
[0104] Step B12, a convolution layer with a convolution kernel size of D×D and a step size of D constitutes the core operation of the block coding layer. This operation divides the features obtained in step B11 into N blocks of size D×D, that is, The channel dimension of each block is C; then the features of these blocks are flattened, and the resulting sequence-level features As the input of the first recursive interaction block. The specific formula is as follows:
[0105] X (0) =PatchEmbed(LeakyReLU(Conv3(I in )))
[0106] Among them, Conv3 represents a 3×3 convolutional layer, LeakyReLU(·) represents the activation function, and PatchEmbed(·) represents the block coding layer operation.
[0107] Furthermore, the step B2 comprises the following steps:
[0108] Step B21, designing recursive interaction blocks, each recursive interaction block includes T parameter-sharing recursive interaction units;
[0109] Step B22: The feature X obtained in step B12 (0) As the input of the first recursive interaction block, after l recursive interaction blocks and T recursions, the output feature is obtained Where l∈{1,2,...,L}. Then, the output of the last recursive interaction block after T recursions is expressed as The specific formula is as follows:
[0110]
[0111] Among them, RIB l represents the lth recursive interaction block, and ° represents the stacking process of the network.
[0112] Furthermore, the step B21 includes the following steps:
[0113] Step B211, designing a recursive interaction unit, which consists of a local block-aware dual attention module, a global context-aware shiftable attention module, an interaction operation between the two modules, and two operations: sequence-level feature reshaping and block-level feature reshaping;
[0114] Step B212: As the input feature of the lth recursive interaction block, the output feature of the current recursive interaction block is obtained after passing through T parameter-sharing recursive interaction units. The specific formula is as follows:
[0115]
[0116] Among them, RIU t represents the t-th recursive interaction unit, t∈{1,2,...,T}, Indicates repeated processing of the network, that is, parameter reuse.
[0117] Further, the step B211 is implemented as follows:
[0118] Step B2111, design a local block-aware dual attention module, the core of which includes a convolution unit, a channel attention unit, and a spatial attention unit. First reshape it into block-level features Then, feature embedding is performed through a convolution unit, which is composed of a 3×3 convolution layer, a ReLu activation function, and a 3×3 convolution layer stacked in sequence. The mapped feature is represented as The specific formula is as follows:
[0119]
[0120]
[0121] Among them, Reshape1(·) is a reshaping operation that converts the N×C sequence-level features into The block-level features of , ConvUnit(·) represents the convolution unit;
[0122] Step B2112: Change the features of step B2211 The channel attention unit and the spatial attention unit are sent to obtain the weighted features of the channel attention and spatial attention respectively, and then the two features are concatenated, and a convolution layer composed of a convolution kernel of size 1×1 is used to map the features to a low-dimensional space, and finally Perform residual connection operation to obtain the output features of the local block-aware dual attention module The specific formula is as follows:
[0123]
[0124] Among them, CA represents the channel attention unit, SA represents the spatial attention unit, [·;·] represents the feature concatenation operation, and Conv1 represents the 1×1 convolutional layer;
[0125] Step B2113: Design the interactive operation from the local block-aware dual attention module to the global context-aware shiftable attention module. For the features output by the local block-aware dual attention module Reshape it back to sequence-level features and then compare it with the features Fusion as input to the global context-aware shiftable attention module The specific formula can be expressed as:
[0126]
[0127] Reshape1(·) is a reshaping operation. The conversion of N×C sequence-level features into block-level features;
[0128] Step B2114, design a global context-aware shiftable attention module, the core of which includes layer normalization operation, shiftable multi-head self-attention and multi-layer perceptron. The features are then sent to the normalization layer and the offset multi-head self-attention layer, and the obtained features are then combined with Residual connection, output intermediate features Then, The features are sent to the normalization layer and the multi-head perceptron in turn, and then Residual connection, outputting the final features of the global context-aware shiftable attention module The specific formula is as follows:
[0129]
[0130]
[0131] Among them, LN(·) represents the layer normalization operation, MHSA represents the shiftable multi-head self-attention, and MLP represents the multi-head perceptron;
[0132] Step B2115: Design the interactive operation of the global context-aware shiftable attention module to the local block-aware dual attention module. For the features output by the global context-aware shiftable attention module Reshape it back to block-level features and then compare it with the features The output of the dual attention module is fused to update the local patch-aware The specific formula can be expressed as:
[0133]
[0134] Among them, Reshape1(·) means reshaping the sequence-level features into block-level features.
[0135] Further, the step B2112 is implemented as follows:
[0136] Step B21121: Design the channel attention unit. First, input the feature The features are compressed into a global average pooling layer to obtain channel descriptors. Next, two 1×1 convolutional layers are used with a ReLu activation function in between to perform dimensionality reduction and dimensionality increase, respectively. Finally, the channel weights are limited between 0 and 1 through the Sigmoid activation function to obtain the channel attention weights. Finally, through the multiplication operation With channel weight α c Multiply them together to get the channel attention weighted features The specific formula can be expressed as:
[0137]
[0138]
[0139] Among them, AvgPool(·) represents the global average pooling operation, Conv1 represents the 1×1 convolutional layer, ReLu(·) is the activation function, σ(·) is the Sigmoid function, represents element-wise multiplication;
[0140] Step B21122: Design the spatial attention unit. First, input features Perform global maximum pooling and global average pooling operations along the channel dimension to obtain the maximum pooling feature map and the average pooling feature map, and then concatenate them along the channel dimension. Next, use a 3×3 convolution layer to map the feature map back to a 1-dimensional feature with a channel, and then use the Sigmoid activation function to limit the weight between 0 and 1 to obtain the spatial attention weight. Finally, through the multiplication operation With the spatial attention weight α s Multiply them together to get the spatial attention weighted features The specific formula can be expressed as:
[0141]
[0142]
[0143] Among them, AvgPool(·) represents the global average pooling operation, MaxPool(·) represents the global maximum pooling operation, [·;·] represents the concatenation operation, Conv3 represents the 3×3 convolutional layer, σ(·) is the Sigmoid function, Represents element-wise multiplication.
[0144] Further, step B3 is implemented as follows:
[0145] Step B31, the core operation of the block decoding layer is composed of a bilinear upsampling layer and a 3×3 convolutional layer. This operation converts the N×C size feature of the last recursive interaction block output obtained in step B22 into Reshape to The block-level features of size are then upsampled to H×W×C features by a bilinear upsampling layer, and finally mapped by a 3×3 convolutional layer to obtain the features with reduced dimensionality.
[0146] Step B32: For the features output from step B31, a 3×3 convolutional layer is used to map them back to the image space and then combined with the input features I in The residual connection obtains the enhanced image I en The specific formula is as follows:
[0147]
[0148] Among them, PatchUnEmbed(·) is a block decoding operation, and Conv3 represents a 3×3 convolutional layer.
[0149] Further, step C is implemented as follows:
[0150] Step C: Design the loss function, using Smooth L1 loss and VGG perceptual loss l perceptualComposition, the total objective loss function of the network l total It is expressed as follows:
[0151]
[0152]
[0153] l perceptual =||(VGG 3,8,15 (I en )-VGG 3,8,15 (I gt ))|| 2
[0154] Where λ is a balance parameter, E=I en -I gt , I gt is an image with normal illumination. ||·|| 2 Indicates the calculation of mean square error, VGG 3,8,15 (·) indicates that the features of layer 3, layer 8, and layer 15 are extracted using the VGG-16 classification model pre-trained on the ImageNet dataset.
[0155] Further, the step D is implemented as follows:
[0156] Step D1, randomly divide the training data set obtained in step A into several batches, each batch containing N pairs of images;
[0157] Step D2: input low-light image I in , after the recursive interactive attention enhancement network in step B, the enhanced image I is obtained en , use the formula in step C to calculate the loss l total ;
[0158] Step D3: Calculate the gradient of the parameters in the network using the back propagation method according to the loss, and update the network parameters using the Adam optimization method;
[0159] Step D4: Repeat steps D1 to D3 in batches to obtain a recursive interactive attention enhancement network model.
[0160] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.
Claims
1. A low-light image enhancement method based on recursive interactive attention, It is characterized in that The steps include: Step A: preprocess the input image, including image pairing, cropping, and data enhancement, to obtain a training data set; Step B: designing a recursive interactive attention enhancement network, which consists of an input mapping module, a recursive interactive attention enhancement network and an output mapping module; Step C, designing a loss function for training the network designed in step B; Step D: training the recursive interactive attention enhancement network using the training dataset; Step E: input the image to be tested into the network, and use the trained network to generate a normal illumination image; The specific implementation steps of step B are as follows: Step B1, designing an input mapping module, including a convolution layer, an activation function and a block coding layer, for extracting features from the input low-light image, dividing it into blocks and converting it into sequence-level features; Step B2: Design a recursive interactive attention enhancement network, which consists of L recursive interaction blocks whose parameters are not shared, to achieve deep enhancement of features; Step B3, design an output mapping module, including a block decoding layer and a convolutional layer, to reorganize the output features of the recursive interactive attention enhancement network into complete image features, and project the features into an image space with 3 channels; The specific implementation steps of step B1 are as follows: Step B11: A convolutional layer is formed by a convolution kernel of size 3×3, and LeakyReLU is used as the activation function to transform the input image Perform feature extraction, where H and W are the height and width of the low-light image respectively; Step B12, a convolution layer with a convolution kernel size of D×D and a step size of D constitutes the core operation of the block coding layer. This operation divides the features obtained in step B11 into N blocks of size D×D, that is, The channel dimension of each block is C; then the features of these blocks are flattened, and the resulting sequence-level features As the input of the first recursive interaction block; the specific formula is as follows: X (0) =PatchEmbed(LeakyReLU(Conv3(I in ))) Among them, Conv3 represents a 3×3 convolutional layer, LeakyReLU(·) represents an activation function, and PatchEmbed(·) represents a block coding layer operation; The specific implementation steps of step B2 are as follows: Step B21, designing recursive interaction blocks, each recursive interaction block includes T parameter-sharing recursive interaction units; Step B22: The feature X obtained in step B12 (0) As the input of the first recursive interaction block, after l recursive interaction blocks and T′ recursions, the output feature is obtained where l∈{1,2,...,L}; then, the output of the last recursive interaction block after T′ recursions is expressed as The specific formula is as follows: Among them, RIB l represents the lth recursive interaction block, Indicates the stacking processing of the network; The specific implementation steps of step B21 are as follows: Step B211, designing a recursive interaction unit, which consists of a local block-aware dual attention module, a global context-aware shiftable attention module, an interaction operation between the two modules, and two operations: sequence-level feature reshaping and block-level feature reshaping; Step B212: As the input feature of the lth recursive interaction block, the output feature of the current recursive interaction block is obtained after passing through t parameter-sharing recursive interaction units. The specific formula is as follows: Among them, RIU t represents the t-th recursive interaction unit, t∈{1,2,...,T}, Indicates repeated processing of the network, i.e. parameter reuse; The specific implementation steps of step B211 are as follows: Step B2111, design a local block-aware dual attention module, the core of which includes a convolution unit, a channel attention unit and a spatial attention unit; for the input features First reshape it into block-level features Then, feature embedding is performed through a convolution unit, which is composed of a 3×3 convolution layer, a ReLu activation function, and a 3×3 convolution layer stacked in sequence. The mapped feature is represented as The specific formula is as follows: Among them, Reshape1(·) is a reshaping operation that converts the N×C sequence-level features into The block-level features of , ConvUnit(·) represents the convolution unit; Step B2112: Change the features of step B2211 The channel attention unit and the spatial attention unit are sent to obtain the weighted features of the channel attention and spatial attention respectively, and then the two features are concatenated, and a convolution layer composed of a convolution kernel of size 1×1 is used to map the features to a low-dimensional space, and finally Perform residual connection operation to obtain the output features of the local block-aware dual attention module The specific formula is as follows: Among them, CA represents the channel attention unit, SA represents the spatial attention unit, [·;·] represents the feature concatenation operation, and Conv1 represents the 1×1 convolutional layer; Step B2113: Design the interactive operation from the local block-aware dual attention module to the global context-aware shiftable attention module. For the features output by the local block-aware dual attention module Reshape it back to sequence-level features and then compare it with the features Fusion as input to the global context-aware shiftable attention module The specific formula is: Reshape1(·) is a reshaping operation. The conversion of N×C sequence-level features into block-level features; Step B2114, design a global context-aware shiftable attention module, the core of which includes layer normalization operation, shiftable multi-head self-attention and multi-layer perceptron; for the input features The features are then sent to the normalization layer and the offset multi-head self-attention layer, and the obtained features are then combined with Residual connection, output intermediate features Then, The features are sent to the normalization layer and the multi-head perceptron in turn, and then Residual connection, outputting the final features of the global context-aware shiftable attention module The specific formula is as follows: Among them, LN(·) represents the layer normalization operation, MHSA represents the shiftable multi-head self-attention, and MLP represents the multi-head perceptron; Step B2115: Design the interactive operation of the global context-aware shiftable attention module to the local block-aware dual attention module. For the features output by the global context-aware shiftable attention module Reshape it back to block-level features and then compare it with the features The output of the dual attention module is fused to update the local patch-aware The specific formula is: Among them, Reshape1(·) means reshaping the sequence-level features into block-level features; The specific implementation of step B2112 is: Step B21121: Design the channel attention unit. First, input the feature The features are compressed into a global average pooling layer to obtain channel descriptors. Next, two 1×1 convolutional layers are used with a ReLu activation function in between to reduce and increase the dimensionality. Finally, the channel weights are limited between 0 and 1 through the Sigmoid activation function to obtain the channel attention weights. Finally, through the multiplication operation With channel weight α c Multiply them together to get the channel attention weighted features The specific formula is: Among them, Avgpool(·) represents the global average pooling operation, Conv1 represents the 1×1 convolution layer, ReLu(·) is the activation function, σ(·) is the Sigmoid function, represents element-wise multiplication; Step B21122: Design the spatial attention unit. First, input features Perform global maximum pooling and global average pooling operations along the channel dimension to obtain the maximum pooling feature map and the average pooling feature map, and then concatenate them along the channel dimension; then use a 3×3 convolution layer to map the feature map back to the 1-dimensional feature of the channel, and then use the Sigmoid activation function to limit the weight between 0 and 1 to obtain the spatial attention weight Finally, through the multiplication operation and the spatial attention weight α s Multiply them together to get the spatial attention weighted features The specific formula is: Among them, AvgPool(·) represents the global average pooling operation, MaxPool(·) represents the global maximum pooling operation, [·;·] represents the concatenation operation, Conv3 represents the 3×3 convolutional layer, σ(·) is the Sigmoid function, represents element-wise multiplication; The specific implementation steps of step B3 are as follows: Step B31, the core operation of the block decoding layer is composed of a bilinear upsampling layer and a 3×3 convolutional layer. This operation converts the N×C size feature of the last recursive interaction block output obtained in step B22 into Reshape to The block-level features of size are then upsampled to H×W×C features by a bilinear upsampling layer, and finally mapped by a 3×3 convolutional layer to obtain the features with reduced dimension. Step B32: For the features output from step B31, a 3×3 convolutional layer is used to map them back to the image space and then combined with the input features I in The residual connection obtains the enhanced image I en ; The specific formula is as follows: Among them, PatchUnEmbed(·) is a block decoding operation, and Conv3 represents a 3×3 convolutional layer; The specific implementation of step C is: Step C: Design the loss function, using Smooth L1 loss and VGG perceptual loss l perceptual Composition, the total objective loss function of the network l total It is expressed as follows: l perceptual =||(VGG 3,8,15 (I en )-VGG 3,8,15 (I gt ))|| 2 Where λ is a balance parameter, E = I en -I gt , I gt is an image with normal illumination; ||·|| 2 Indicates the calculation of mean square error, VGG 3,8,15 (·) indicates that the features of layer 3, layer 8, and layer 15 are extracted using the VGG-16 classification model pre-trained on the ImageNet dataset.
2. The low-light image enhancement method based on recursive interactive attention according to claim 1, It is characterized in that The specific implementation steps of step A are as follows: Step A1: pairing the normal illumination image with the low illumination image as a label image; Step A2: randomly crop each low-light image of size H×W×3 into an image of size P×P×3, and use the same random cropping method for its corresponding normal-light image to ensure that they have the same size and position, where H and W are the height and width of the low-light image and the normal-light image, and P is the height and width of the cropped image; Step A3: For each training paired image, randomly apply one of the following eight data augmentation methods: keep the original image, flip vertically, rotate 90 degrees, rotate 90 degrees and then flip vertically, rotate 180 degrees, rotate 180 degrees and then flip vertically, rotate 270 degrees, and rotate 270 degrees and then flip vertically.
3. The low-light image enhancement method based on recursive interactive attention according to claim 1, It is characterized in that The specific implementation steps of step D are as follows: Step D1, randomly divide the training data set obtained in step A into several batches, each batch containing N' pairs of images; Step D2: input low-light image I in , after the recursive interactive attention enhancement network in step B, the enhanced image I is obtained en , use the formula in step C to calculate the loss l total ; Step D3: Calculate the gradient of the parameters in the network using the back propagation method according to the loss, and update the network parameters using the Adam optimization method; Step D4: Repeat steps D1 to D3 in batches to obtain a recursive interactive attention enhancement network model.
Citation Information
Patent Citations
Depth map super-resolution method based on multistage recursion guidance and progressive supervision
CN110111254A
Object image re-identification method based on multi-feature information capture and correlation analysis
WO2023273290A1