Low illumination image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction
Through a low-illumination image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction, combined with the bilateral interaction of local perception and global perception modules, the problems of image details loss and insufficient global consistency in the prior art are solved, and a more natural and detailed image enhancement effect is achieved.
Patent Information
- Application Number
- CN202310859559.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-13
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-07-13
AI Technical Summary
The existing low-illumination image enhancement methods have shortcomings in detail recovery and global consistency representation, resulting in loss of image details and poor global consistency after enhancement.
The low-illumination image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction is adopted. Through the bilateral interaction between the local perception module and the global perception module, combined with local details and global structural information, the natural and detailed enhancement of the image is achieved.
This method can not only effectively restore the detailed information of low-illumination images, but also improve the global consistency representation of the image, making the generated enhanced images more natural and detailed.
Smart Images

Figure CN116703783B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image and video processing and computer vision technology, in particular to a low-illumination image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction. Background Art
[0002] With the continuous advancement of science and technology and the rapid development of camera technology, low-light images are becoming widely popular in daily life and various industries at an alarming rate. At night, indoors or in other dimly lit scenes, we are faced with more and more needs to capture photos or record videos in low-light environments. However, due to the lack of available light in the environment, the camera sensor cannot accurately capture details and generate clear images, resulting in blurry images, excessive noise, color distortion and other problems. Therefore, low-light image enhancement has become an important topic that has received much attention. The goal of low-light image enhancement is to improve the quality of images taken under low-light conditions through specific technologies and algorithms, mainly including hardware improvements and software algorithms. In terms of hardware, camera manufacturers enhance imaging capabilities in low-light conditions by improving sensor sensitivity, increasing pixel size and improving lens quality. These hardware improvements enable cameras to better capture light and provide higher quality image materials. Software algorithms play a key role in post-processing, reducing noise through noise reduction algorithms, improving image visibility through contrast enhancement algorithms, and restoring lost details through image restoration algorithms.
[0003] Low-light image enhancement technology has broad application prospects in various fields. First, in the field of security monitoring, this technology can improve the quality of nighttime surveillance images and help police or security personnel better identify and analyze potential threats. In the past, due to insufficient light, surveillance cameras often could not clearly capture important details, which affected the effect of security monitoring. With the help of low-light image enhancement technology, surveillance images can be significantly improved, subtle movements and behaviors can be presented more clearly, and security prevention capabilities can be effectively improved. Secondly, in the field of medical imaging, low-light image enhancement technology also has great potential. Medical imaging is essential for accurate diagnosis and treatment, and in some cases, such as intraoperative operations or examinations of specific cases, light may be limited. Low-light image enhancement technology provides a feasible solution that enables doctors to observe important details such as tiny lesions and tissue structures, thereby achieving more accurate diagnosis and treatment plans. For example, in tumor surgery, low-light image enhancement technology can help doctors more clearly distinguish the boundaries between cancerous tissue and healthy tissue, reduce surgical risks, and increase the success rate of surgery and the patient's chance of survival. In addition, in the field of autonomous driving, low-light image enhancement plays a vital role in vehicle perception and decision-making. Autonomous driving technology is becoming an important part of future transportation, and in poor lighting conditions, vehicle perception systems need reliable image data to accurately identify roads, obstacles, and other traffic participants. Low-light image enhancement technology can help improve the performance of vehicle perception systems, enabling them to better adapt to driving environments under different lighting conditions. By reducing noise in images, enhancing details, and improving image contrast, this technology can improve the accuracy of driving decisions, thereby increasing driving safety and reducing the risk of accidents.
[0004] Recently, deep learning models have made remarkable progress in the field of low-light image enhancement. They are able to learn powerful and generalizable prior knowledge from large-scale datasets, making them the dominant approach in this field. Currently, there are two common convolutional neural network architecture designs that are widely used: encoder-decoder and high-resolution (single-scale) feature processing. The encoder-decoder model gradually maps the image to a low-resolution representation and inversely maps it back to the original resolution. However, this approach leads to the loss of fine spatial details, making it difficult to restore them in subsequent stages. On the other hand, high-resolution (single-scale) networks avoid downsampling operations, but their receptive fields are limited and the encoding of contextual information is inefficient. Summary of the invention
[0005] In view of this, the purpose of the present invention is to provide a low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction. The method combines local perception and global perception enhancement, can fully consider the local details and global structure of the image, so that the enhanced image is more natural and detailed; through local-global and global-local interactions, the local perception module can provide detail information to the global perception module to help it better understand the contextual relationship of the image, and the global perception module can provide overall brightness information to the local perception module to help it perform detail enhancement. Such a bilateral interaction strategy not only enhances the reconstructed details when enhancing low-light images, but also effectively promotes the representation of global consistency.
[0006] To achieve the above object, the present invention adopts the following technical solution: a low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction, comprising the following steps:
[0007] Step A: preprocessing the input image, including image pairing, cropping and data enhancement, to obtain a training data set;
[0008] Step B: design a low-light image interactive enhancement network based on CNN-Transformer. The low-light image interactive enhancement network based on CNN-Transformer consists of an input mapping module, L interactive enhancement blocks and an output mapping module;
[0009] Step C, designing a loss function for training the network designed in step B;
[0010] Step D: Use the training data set to train the CNN-Transformer based low illumination image interactive enhancement network;
[0011] Step E: input the image to be tested into the network, and use the trained network to generate a normal illumination image.
[0012] In a preferred embodiment, the specific implementation steps of step A are as follows:
[0013] Step A1: Pair the normal illumination image with the low illumination image, wherein the normal illumination image is used as the label image;
[0014] Step A2: randomly crop each low-light image of size H×W×3 into an image of size P×P×3, and use the same random cropping method for its corresponding normal-light image to ensure that they have the same size and position, where H and W are the height and width of the low-light image and the normal-light image, and P is the height and width of the cropped image;
[0015] Step A3: For each training paired image, randomly apply one of the following eight data augmentation methods: keep the original image, flip vertically, rotate 90 degrees, rotate 90 degrees and then flip vertically, rotate 180 degrees, rotate 180 degrees and then flip vertically, rotate 270 degrees, and rotate 270 degrees and then flip vertically.
[0016] In a preferred embodiment, the specific implementation steps of step B are as follows:
[0017] Step B1, design the input mapping module, the core of which contains two convolution units. Each unit consists of a convolution layer with a convolution kernel size of 3×3 and a stride of 1, a ReLu activation function and a bilinear downsampling layer in sequence to achieve low-light input images. The feature extraction of Where H and W are the height and width of the low-light image, respectively, D represents the total downsampling ratio, and C represents the number of channels for feature extraction;
[0018] Step B2: Design an interactive enhancement block, which consists of a local perception enhancement module, a global perception enhancement module, a local-to-global interactive operation, and a global-to-local interactive fusion module. The low-light image is enhanced by hybrid modeling and bilateral interaction. For the feature X extracted in step B1, in , the enhanced feature representation after L interactive enhancement blocks is The available formulas are described as:
[0019]
[0020] Among them, IEB stands for Interaction Enhancement Block, Represents the stacking process of interactive enhancement blocks;
[0021] Step B3, design the output mapping module, the core of which contains two convolution units, each of which is composed of a bilinear upsampling layer, a convolution layer with a convolution kernel size of 3×3 and a step size of 1, and a ReLu activation function in sequence; for the feature X extracted in step B1 in , connected with the output feature X of step B2 through residual connection out Add as the input of this module and reconstruct into a 3-channel image
[0022] Step B4: Features reconstructed in step B3 With input feature I in By adding them together through residual connection, the enhanced 3-channel image is obtained, which is expressed as
[0023] In a preferred embodiment, the specific implementation steps of step B2 are as follows:
[0024] Step B21: Input the feature X extracted by the mapping module in step B1 in , before being sent to the interactive enhancement block, two copies are first copied as the input of the local perception enhancement module and the global perception enhancement module respectively; here the input of the local perception enhancement module in the lth interactive enhancement block is represented as The input of the global perception enhancement module is expressed as
[0025] Step B22, design a local perception enhancement module, which is mainly composed of channel context modeling, channel transformation, spatial context modeling and spatial transformation units; for the input feature D l-1 , the channel correction factor Z is obtained by channel context modeling and channel transformation l ; The spatial correction factor S is obtained by spatial context modeling and spatial transformation l ; Then, the channel correction factor Z l and the spatial correction factor S l Acting on the input feature D l-1 , to obtain the output feature D of the local perception enhancement module l ; The process is expressed by the formula:
[0026]
[0027] in, represents bitwise multiplication, Indicates bitwise addition;
[0028] Step B23: Design local to global interactive operations for the output D of the local perception enhancement module. l , and input it into the global perception enhancement module as the search domain feature;
[0029] Step B24, design a global perception enhancement module, which is mainly composed of a spatial cross attention layer, a normalization layer, a channel self-attention layer, and a feedforward network; for the input feature G l-1 , and use it as the query domain feature, together with the search domain feature D in step B23 l Together, they are first sent to the spatial cross attention and normalization layers in sequence, and the obtained features are then combined with G l-1 Residual connection, output intermediate features Then, The features are sent to the channel self-attention layer and normalization layer in turn, and then Residual connection, output intermediate features Finally, Send it to the feedforward network to get the output feature G of the global perception enhancement module l ; The process is expressed by the formula:
[0030]
[0031]
[0032]
[0033] Among them, LN(·) represents the layer normalization operation, SMCA represents spatial cross attention, CMSA represents channel self-attention, and FFN represents feedforward network;
[0034] Step B25, design a global to local interactive fusion module, which reuses the global perception enhancement module structure; for the output feature D of the local perception enhancement module l And the output module G of the global perception enhancement module l , respectively, and use them as the query domain and search domain, and send them to the spatial cross attention and normalization layers in turn. The obtained features are then combined with D l Residual connection, output intermediate features Then, The features are sent to the channel self-attention layer and normalization layer in turn, and then Residual connection, output intermediate features Finally, Send it to the feedforward network to get the final output feature X l , which is the output feature of the lth interaction enhancement block; the process is expressed by the formula:
[0035]
[0036]
[0037]
[0038] In a preferred embodiment, the specific implementation steps of step B22 are as follows:
[0039] Step B221: For input features First, it is processed by the channel context modeling unit; specifically, the features are first compressed by global average pooling to obtain the channel correction factor Next, the correction factor is concatenated with the historical l-1 channel correction factors to obtain the total correction factor representation It can be expressed as:
[0040]
[0041]
[0042] Among them, AvgPool(·) represents global average pooling, [·;·] represents the concatenation operation;
[0043] Step B222: Obtain the total correction factor for step B221 The channel transformation unit continues processing; it passes through a 1×1 convolution layer, a ReLu activation function, a 1×1 convolution layer, and a Sigmoid function in turn to obtain the current channel correction factor The process is expressed in the formula:
[0044]
[0045] Among them, Cov1(·) represents a 1×1 convolutional layer, ReLu(·) represents a ReLu activation function, and Sigmoid(·) represents a Sigmoid function;
[0046] Step B223: For input features It is also processed by the spatial context modeling unit; specifically, for D l-1 , firstly, the global average pooling and global maximum pooling operations along the channel dimension are performed to obtain a size of The global average feature map and the global maximum feature map are obtained, and the two feature maps are concatenated along the channel dimension; the process is expressed as follows:
[0047]
[0048] Among them, Avgpool(·) and MaxPool(·) represent global average pooling and global maximum pooling operations respectively. Represents the feature map after splicing;
[0049] Step B224: feature map of step B223 The spatial transformation unit continues to process the image, which passes through a 1×1 group convolution layer, a 3×3 convolution layer, and the Tanh function to obtain the spatial correction factor. The formula is:
[0050]
[0051] Among them, GroupConv(·) represents a 1×1 group convolution layer, conv3(·) represents a 3×3 convolution layer, and Tanh(·) represents the Tanh function;
[0052] Step B225: For input features First, compare it with the channel correction factor obtained in step B222 Multiply, and then add the spatial correction factor obtained in step B224 Add together to get the output D of the local perception enhancement module l .
[0053] In a preferred embodiment, the specific implementation steps of step B24 are as follows:
[0054] Step B241, design a spatial cross attention layer, for the input feature G l-1 and D l , first reshaped into features of size C×N, where Then three groups of 1×1 convolutions are used to transform G l-1 Mapped to the "query" matrix Q sg , D l Mapped to the "key" matrix K sd and the "value" matrix V sd ; Further, the three matrices are each divided into h heads, each head has a size of N and a channel number So we get the "query" matrix of the i-th head The "key" matrix and the "value" matrix where i∈[1,h];
[0055] Step B242: K sd,i Calculate softmax along the spatial dimension to get the weight of the spatial dimension Q sg,i Calculate softmax along the channel dimension to get the weight of the channel dimension Then CK sd,i Perform transpose operation and V sd,i Perform matrix multiplication to obtain the global information of N×N space, and then add it to SQ sg,i Multiply by the aggregated channel weight; finally, concatenate the features aggregated by all attention heads and perform a linear transformation with a 1×1 convolution to obtain the features enhanced by the spatial cross attention layer The process is expressed as:
[0056] CK sd,i =softmax1(K sd,i )
[0057] SQ sg,i =softmax2(Q sg,i )
[0058]
[0059]
[0060] Among them, softmax1(·) represents the softmax function calculated along the spatial dimension, softmax2(·) represents the softmax function calculated along the channel dimension, [·;·] represents the concatenation operation, and W o Represents the learnable parameter matrix of 1×1 convolution;
[0061] Step B243: Enhanced features of the spatial cross attention layer in step B242 After transposing it, it is sent to the normalization layer, and the obtained features are then combined with G l-1 Residual connection, the obtained features As the input of the next stage; expressed by the formula:
[0062]
[0063] Among them, Transpose(·) means transposing the C×N size feature to the N×C feature;
[0064] Step B244, design the channel self-attention layer, for the input features Use three sets of 1×1 convolutions to map it into a “query” matrix Q cg , the "key" matrix K cg and the "value" matrix V cg ; Split each of the three matrices into h heads, each head has a size of N and a number of channels So we get the "query" matrix of the i-th head The "key" matrix and the "value" matrix where i∈[1,h];
[0065] Step B245: K cg,i Calculate softmax along the channel dimension to get the weight of the channel dimension Q cg,i Calculate softmax along the spatial dimension to get the weight of the spatial dimension Then CK cg,i Perform transpose operation and V sd,i Perform matrix multiplication to get a matrix of size Channel global information, then with SQ cg,i Multiply by the aggregated spatial weight; finally, concatenate the features aggregated by all attention heads and perform a linear transformation with a 1×1 convolution to obtain the features enhanced by the channel self-attention layer The process is expressed in the formula:
[0066] CK cg,i =softmax2(K cg,i )
[0067] SQ cg,i =softmax1(Q sg,i )
[0068]
[0069]
[0070] Among them, softmax2(·) represents the softmax function calculated along the channel dimension, softmax1(·) represents the softmax function calculated along the spatial dimension, [·;·] represents the concatenation operation, and W p Represents the learnable parameter matrix of 1×1 convolution;
[0071] Step B246: Enhanced features of the channel cross attention layer in step B245 Send it to the normalization layer, and then get the features Residual connection to obtain features
[0072] Step B247: Send it to the feedforward network to get the output feature G of the global perception enhancement module l .
[0073] In a preferred embodiment, the specific implementation steps of step B247 are as follows:
[0074] Step B2471: Features enhanced by channel self-attention First, the channel is doubled by a 1×1 convolution, and then a 3×3 depth-separable convolution is used for feature mapping; then, the mapped features are divided into two parts along the channel dimension, represented as and The process is expressed as:
[0075]
[0076] Among them, Conv1(·) represents a 1×1 convolutional layer, DWConv3(·) represents a 3×3 depth-separable convolutional layer, and chunk(·) means dividing the features into two parts along the channel dimension;
[0077] Step B2472: and The two features are activated by ReLu and multiplied by each other’s original features to achieve information interaction. Then, a 1×1 convolution is used to transform the features to obtain the output feature G of the global perception enhancement module. l ; The process can be expressed as:
[0078]
[0079] Where ReLu(·) represents the ReLu activation function, W t Represents the learnable parameters of the 1×1 convolution.
[0080] In a preferred embodiment, the specific implementation of step C is as follows:
[0081] Step C: Design the loss function, using L1 loss Structural loss function and VGG perceptual loss Composition, the total objective loss function of the network It is expressed as follows:
[0082]
[0083]
[0084]
[0085]
[0086] Among them, λ 1 , 2 and λ 3 is the equilibrium parameter, I en For enhanced images, I gt is the image with normal illumination; μ en and μ gt is the mean of the two images, σ en and σ gt represents the variance of the two images, c 1 and c 2 are two constant values to prevent the denominator from being 0; ||·|| 2 Indicates the calculation of mean square error, VGG 3,8,15 (·) indicates that the features of layer 3, layer 8, and layer 15 are extracted using the VGG-16 classification model pre-trained on the ImageNet dataset.
[0087] In a preferred embodiment, the specific implementation steps of step D are as follows:
[0088] Step D1, randomly divide the training data set obtained in step A into several batches, each batch containing N pairs of images;
[0089] Step D2: input low-light image I in After the interactive enhancement network in step B, the enhanced image I is obtained en , use the formula in step C to calculate the loss
[0090] Step D3: Calculate the gradient of the parameters in the network using the back propagation method according to the loss, and update the network parameters using the Adam optimization method;
[0091] Step D4: Repeat steps D1 to D3 in batches to obtain an interactive enhancement network model.
[0092] In a preferred embodiment, the specific implementation steps of step E are as follows:
[0093] Step E1: Input a low-light image I of size H×W×3 in Reshape the image into a size of P×P×3, where H and W are the height and width of the low-light image, respectively, and P is the height and width of the reshaped image;
[0094] Step E2: Send the reshaped image to the interactive enhancement network in step B to obtain an enhanced image of size P×P×3; Step E3: Reshape the enhanced image of size P×P×3 into an enhanced image I of size H×W×3 en .
[0095] Compared with the prior art, the present invention has the following beneficial effects: first, the present invention encodes and downsamples the features in the input mapping module, which can effectively reduce the complexity of subsequent calculations while extracting multi-scale features. Secondly, unlike the traditional Transformer that implements block coding based on linear layers and non-overlapping embedding layers, the present invention adopts an overlapping embedding method, which can simultaneously retain position and adjacent feature information. Furthermore, the present invention designs an interactive enhancement network, including a local perception module and a global perception module, which enhance the details and texture of the image by using convolution-based spatial and channel corrections, and better understand the overall content and semantic information of the image by using spatial cross attention and channel self-attention. In addition, the present invention designs interactive operations to promote global consistent representation. Finally, the present invention uses an output mapping module to convert the enhanced feature map into a final image, which can better retain the detailed information of the image and improve the quality and brightness of the image. Unlike other recent low-light image enhancement methods based on Transformer, the present invention not only combines the advantages of local perception and global perception, but also uses interactive operations to effectively improve the learning efficiency of the Transformer structure, and can achieve efficient low-light image enhancement under limited computing resources and storage space, thereby reducing the implementation cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] Figure 1 It is a flow chart of the implementation of the method in the preferred embodiment of the present invention.
[0097] Figure 2 It is a structural diagram of an interactive enhancement network in a preferred embodiment of the present invention.
[0098] Figure 3 It is a structural diagram of the interactive enhancement block in a preferred embodiment of the present invention.
[0099] Figure 4 It is a structural diagram of the local perception module in the preferred embodiment of the present invention. DETAILED DESCRIPTION
[0100] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0101] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0102] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.
[0103] The present invention provides a low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction, such as Figure 1-4 As shown, the following steps are included:
[0104] Step A: preprocessing the input image, including image pairing, cropping and data enhancement, to obtain a training data set;
[0105] Step B: Design a low-light image interactive enhancement network based on CNN-Transformer, which consists of an input mapping module, L interactive enhancement blocks and an output mapping module;
[0106] Step C, designing a loss function for training the network designed in step B;
[0107] Step D: Use the training data set to train the CNN-Transformer based low illumination image interactive enhancement network;
[0108] Step E: input the image to be tested into the network, and use the trained network to generate a normal illumination image.
[0109] Furthermore, the step A comprises the following steps:
[0110] Step A1: Pair the normal illumination image with the low illumination image, wherein the normal illumination image is used as the label image;
[0111] Step A2: randomly crop each low-light image of size H×W×3 into an image of size P×P×3, and use the same random cropping method for its corresponding normal-light image to ensure that they have the same size and position, where H and W are the height and width of the low-light image and the normal-light image, and P is the height and width of the cropped image;
[0112] Step A3: For each training paired image, randomly apply one of the following eight data augmentation methods: keep the original image, flip vertically, rotate 90 degrees, rotate 90 degrees and then flip vertically, rotate 180 degrees, rotate 180 degrees and then flip vertically, rotate 270 degrees, and rotate 270 degrees and then flip vertically.
[0113] Furthermore, the step B comprises the following steps:
[0114] Step B1, design the input mapping module, the core of which contains two convolution units. Each unit consists of a convolution layer with a convolution kernel size of 3×3 and a stride of 1, a ReLu activation function and a bilinear downsampling layer in sequence to achieve low-light input images. The feature extraction of Where H and W are the height and width of the low-light image, respectively, D represents the total downsampling ratio, and C represents the number of channels for feature extraction;
[0115] Step B2: Design an interactive enhancement block, which consists of a local perception enhancement module, a global perception enhancement module, a local-to-global interactive operation, and a global-to-local interactive fusion module. The low-light image is enhanced by hybrid modeling and bilateral interaction. in , the enhanced feature representation after L interactive enhancement blocks is The available formulas are described as:
[0116]
[0117] Among them, IEB stands for Interaction Enhancement Block, Represents the stacking process of interactive enhancement blocks;
[0118] Step B3, design the output mapping module, the core of which contains two convolution units, each of which is composed of a bilinear upsampling layer, a convolution layer with a convolution kernel size of 3×3 and a step size of 1, and a ReLu activation function. in , connected with the output feature X of step B2 through residual connection outAdd as the input of this module and reconstruct into a 3-channel image
[0119] Step B4: Features reconstructed in step B3 With input feature I in By adding them together through residual connection, the enhanced 3-channel image is obtained, which is expressed as
[0120] Furthermore, the step B2 comprises the following steps:
[0121] Step B21: Input the feature X extracted by the mapping module in step B1 in , before being sent to the interactive enhancement block, two copies are first made as the input of the local perception enhancement module and the global perception enhancement module. Here, the input of the local perception enhancement module in the lth interactive enhancement block is represented as The input of the global perception enhancement module is expressed as
[0122] Step B22: Design a local perception enhancement module, which is mainly composed of channel context modeling, channel transformation, spatial context modeling, and spatial transformation. l-1 , the channel correction factor Z is obtained by channel context modeling and channel transformation l ; The spatial correction factor S is obtained by spatial context modeling and spatial transformation l Then, the channel correction factor Z is l and the spatial correction factor S l Acting on the input feature D l-1 , to obtain the output feature D of the local perception enhancement module l The process is expressed in the formula:
[0123]
[0124] in, represents bitwise multiplication, Indicates bitwise addition;
[0125] Step B23: Design local to global interactive operations for the output D of the local perception enhancement module. l , and input it into the global perception enhancement module as the search domain feature;
[0126] Step B24, design a global perception enhancement module, which is mainly composed of a spatial cross attention layer, a normalization layer, a channel self-attention layer, and a feedforward network. l-1 , and use it as the query domain feature, together with the search domain feature D in step B23 lTogether, they are first sent to the spatial cross attention and normalization layers in sequence, and the obtained features are then combined with G l-1 Residual connection, output intermediate features Then, The features are sent to the channel self-attention layer and normalization layer in turn, and then Residual connection, output intermediate features Finally, Send it to the feedforward network to get the output feature G of the global perception enhancement module l The process is expressed in the formula:
[0127]
[0128]
[0129]
[0130] Among them, LN(·) represents the layer normalization operation, SMCA represents spatial cross attention, CMSA represents channel self-attention, and FFN represents feedforward network;
[0131] Step B25: Design a global-to-local interactive fusion module, which reuses the structure of the global perception enhancement module. l And the output module G of the global perception enhancement module l , respectively, and use them as the query domain and search domain, and send them to the spatial cross attention and normalization layers in turn. The obtained features are then combined with D l Residual connection, output intermediate features Then, The features are sent to the channel self-attention layer and normalization layer in turn, and then Residual connection, output intermediate features Finally, Send it to the feedforward network to get the final output feature X l , which is the output feature of the lth interaction enhancement block. The process is expressed as:
[0132]
[0133]
[0134]
[0135] Furthermore, the step B22 includes the following steps:
[0136] Step B221: For input features First, it is processed by the channel context modeling unit. Specifically, the features are first compressed by global average pooling to obtain the channel correction factor Next, the correction factor is concatenated with the historical l-1 channel correction factors to obtain the total correction factor representation It can be expressed as:
[0137]
[0138]
[0139] Among them, AvgPool(·) represents global average pooling, [·;·] represents the concatenation operation;
[0140] Step B222: Obtain the total correction factor for step B221 The channel transformation unit continues processing. It passes through a 1×1 convolution layer, a ReLu activation function, a 1×1 convolution layer, and a Sigmoid function to obtain the current channel correction factor. The process is expressed as:
[0141]
[0142] Where Conv1(·) represents a 1×1 convolutional layer, ReLu(·) represents a ReLu activation function, and Sigmoid(·) represents a Sigmoid function.
[0143] Step B223: For input features It is also processed by the spatial context modeling unit. Specifically, for D l-1 , firstly, the global average pooling and global maximum pooling operations along the channel dimension are performed to obtain a size of The global average feature map and the global maximum feature map are obtained, and the two feature maps are concatenated along the channel dimension. The process is expressed as:
[0144]
[0145] Among them, AvgPool(·) and MaxPool(·) represent global average pooling and global maximum pooling operations respectively. Represents the feature map after splicing;
[0146] Step B224: feature map of step B223 The spatial transformation unit continues to process it. It passes through a 1×1 group convolution layer, a 3×3 convolution layer, and the Tanh function to obtain the spatial correction factor The formula is:
[0147]
[0148] Among them, GroupConv(·) represents a 1×1 group convolution layer, conv3(·) represents a 3×3 convolution layer, and Tanh(·) represents the Tanh function;
[0149] Step B225: For input features First, compare it with the channel correction factor obtained in step B222 Multiply, and then add the spatial correction factor obtained in step B224 Add together to get the output D of the local perception enhancement module l .
[0150] Further, the step B24 includes the following steps:
[0151] Step B241, design a spatial cross attention layer, for the input feature G l-1 and D l , first reshaped into features of size C×N, where Then three groups of 1×1 convolutions are used to transform G l-1 Mapped to the "query" matrix Q sg , D l Mapped to the "key" matrix K sd and the "value" matrix V sd . Further, the three matrices are each divided into h heads, each head has a size of N and a channel number of So we can get the "query" matrix of the i-th head The "key" matrix and the "value" matrix where i∈[1,h];
[0152] Step B242: K sd,i Calculate softmax along the spatial dimension to get the weight of the spatial dimension Q sg,i Calculate softmax along the channel dimension to get the weight of the channel dimension Then CK sd,i Perform transpose operation and V sd,i Perform matrix multiplication to obtain the global information of N×N space, and then add it to SQ sg,i Multiply by the aggregated channel weight. Finally, concatenate the features aggregated by all attention heads and perform a linear transformation with a 1×1 convolution to obtain the features enhanced by the spatial cross attention layer. The process is expressed in the formula:
[0153] CK sd,i=softmax1(K sd,i )
[0154] SQ sg,i =softmax2(Q sg,i )
[0155]
[0156]
[0157] Among them, softmax1(·) represents the softmax function calculated along the spatial dimension, softmax2(·) represents the softmax function calculated along the channel dimension, [·;·] represents the concatenation operation, and W o Represents the learnable parameter matrix of 1×1 convolution;
[0158] Step B243: Enhanced features of the spatial cross attention layer in step B242 After transposing it, it is sent to the normalization layer, and the obtained features are then combined with G l-1 Residual connection, the obtained features As the input of the next stage. It can be expressed as:
[0159]
[0160] Among them, Transpose(·) means transposing the C×N size feature to the N×C feature;
[0161] Step B244, design the channel self-attention layer, for the input features Use three sets of 1×1 convolutions to map it into a “query” matrix Q cg , the "key" matrix K cg and the "value" matrix V cg . Further, the three matrices are each divided into h heads, each head has a size of N and a channel number of So we can get the "query" matrix of the i-th head The "key" matrix and the "value" matrix where i∈[1,h];
[0162] Step B245: K cg,i Calculate softmax along the channel dimension to get the weight of the channel dimension Q cg,i Calculate softmax along the spatial dimension to get the weight of the spatial dimension Then CK cg,i Perform transpose operation and V sd,i Perform matrix multiplication to get a matrix of size Channel global information, then with SQ cg,i Multiply by the aggregated spatial weight. Finally, concatenate the features aggregated by all attention heads and perform a linear transformation with a 1×1 convolution to obtain the features enhanced by the channel self-attention layer. The process is expressed as:
[0163] CK cg,u =softmax2(K cg,u )
[0164] SQ cg,i =softmax1(Q sg,i )
[0165]
[0166]
[0167] Among them, softmax2(·) represents the softmax function calculated along the channel dimension, softmax1(·) represents the softmax function calculated along the spatial dimension, [·;·] represents the concatenation operation, and W p Represents the learnable parameter matrix of 1×1 convolution;
[0168] Step B246: Enhanced features of the channel cross attention layer in step B245 Send it to the normalization layer, and then get the features Residual connection to obtain features
[0169] Step B247: Send it to the feedforward network to get the output feature G of the global perception enhancement module l .
[0170] Further, the step B247 is implemented as follows:
[0171] Step B2471: Features enhanced by channel self-attention First, a 1×1 convolution is used to double the channel, and then a 3×3 depth-separable convolution is used for feature mapping. Next, the mapped features are divided into two parts along the channel dimension, represented as and The process can be expressed as:
[0172]
[0173] Among them, Conv1(·) represents a 1×1 convolutional layer, DWConv3(·) represents a 3×3 depth-separable convolutional layer, and chunk(·) means dividing the features into two parts along the channel dimension;
[0174] Step B2472: and The two features are activated by ReLu and multiplied by each other’s original features to achieve information interaction. Then, a 1×1 convolution is used to transform the features to obtain the output feature G of the global perception enhancement module. l . The process can be expressed as:
[0175]
[0176] Where ReLu(·) represents the ReLu activation function, W t Represents the learnable parameters of the 1×1 convolution.
[0177] Further, step C is implemented as follows:
[0178] Step C: Design the loss function, using L1 loss Structural loss function and VGG perceptual loss Composition, the total objective loss function of the network It is expressed as follows:
[0179]
[0180]
[0181]
[0182]
[0183] Among them, λ 1 , 2 and λ 3 is the equilibrium parameter, I en For enhanced images, I gt is an image with normal illumination. en and μ gt is the mean of the two images, σ en and σ gt represents the variance of the two images, c 1 and c 2 are two constant values to prevent the denominator from being zero. ||·|| 2 Indicates the calculation of mean square error, V GG3,8,15 (·) indicates that the features of layer 3, layer 8, and layer 15 are extracted using the VGG-16 classification model pre-trained on the ImageNet dataset.
[0184] Further, the step D is implemented as follows:
[0185] Step D1, randomly divide the training data set obtained in step A into several batches, each batch containing N pairs of images;
[0186] Step D2: input low-light image I in After the interactive enhancement network in step B, the enhanced image I is obtained en , use the formula in step C to calculate the loss
[0187] Step D3: Calculate the gradient of the parameters in the network using the back propagation method according to the loss, and update the network parameters using the Adam optimization method;
[0188] Step D4: Repeat steps D1 to D3 in batches to obtain an interactive enhancement network model.
[0189] Further, the step E is implemented as follows:
[0190] Step E1: Input a low-light image I of size H×W×3 in Reshape the image into a size of P×P×3, where H and W are the height and width of the low-light image, respectively, and P is the height and width of the reshaped image;
[0191] Step E2, sending the reshaped image to the interactive enhancement network in step B to obtain an enhanced image of size P×P×3;
[0192] Step E3: Reshape the enhanced image of size P×P×3 into an enhanced image I of size H×W×3 en .
[0193] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.
[0194] The present invention aims to further solve the problems of loss of enhanced image details and poor global consistency caused by the use of pure convolutional neural network structure in low-light image enhancement methods. To this end, a low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction is designed. First, an input mapping module is designed to encode features; then a low-light image interactive enhancement network based on CNN-Transformer is designed, including designing a local perception enhancement module, and repairing degraded features through channel and spatial correction units, designing a global perception enhancement module to perceive the spatial features and channel features of the image in a global range, better suppress artifacts, and design interactive operations to promote global consistency representation to obtain more accurate image enhancement effects; finally, an output mapping module is designed to convert the enhanced feature map into the final image. This low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction combines the advantages of convolutional neural networks and Transformer models: the Transformer model has a powerful attention mechanism, can capture global context information, and is suitable for long-distance dependency modeling. By introducing the Transformer component, the network's perception ability can be enhanced, enabling it to better use global information for reasoning; at the same time, it retains the advantages of convolution operations, can effectively process local features, and has high computational efficiency. Therefore, the CNN-Transformer hybrid modeling method can better solve the problems of detail loss and global consistency.
Claims
1. Low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction, It is characterized in that The steps include: Step A: preprocessing the input image, including image pairing, cropping and data enhancement, to obtain a training data set; Step B: design a low-light image interactive enhancement network based on CNN-Transformer. The low-light image interactive enhancement network based on CNN-Transformer consists of an input mapping module, L interactive enhancement blocks and an output mapping module; Step C, designing a loss function for training the network designed in step B; Step D: Use the training data set to train the CNN-Transformer based low illumination image interactive enhancement network; Step E: input the image to be tested into the network, and use the trained network to generate a normal illumination image; The specific implementation steps of step B are as follows: Step B1, design the input mapping module, the core of which contains two convolution units. Each unit consists of a convolution layer with a convolution kernel size of 3×3 and a stride of 1, a ReLu activation function and a bilinear downsampling layer in sequence to achieve low-light input images. The feature extraction of Where H and W are the height and width of the low-light image, respectively, D represents the total downsampling ratio, and C represents the number of channels for feature extraction; Step B2: Design an interactive enhancement block, which consists of a local perception enhancement module, a global perception enhancement module, a local-to-global interactive operation, and a global-to-local interactive fusion module. The low-light image is enhanced by hybrid modeling and bilateral interaction. For the feature X extracted in step B1, in , the enhanced feature representation after L interactive enhancement blocks is The available formulas are described as: Among them, IEB stands for Interaction Enhancement Block, Represents the stacking process of interactive enhancement blocks; Step B3, design the output mapping module, the core of which contains two convolution units, each of which is composed of a bilinear upsampling layer, a convolution layer with a convolution kernel size of 3×3 and a step size of 1, and a ReLu activation function in sequence; for the feature X extracted in step B1 in , connected with the output feature X of step B2 through residual connection out Add as the input of this module and reconstruct into a 3-channel image Step B4: Features reconstructed in step B3 With input feature I in By adding them together through residual connection, the enhanced 3-channel image is obtained, which is expressed as 2. The low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction according to claim 1, It is characterized in that The specific implementation steps of step A are as follows: Step A1: Pair the normal illumination image with the low illumination image, wherein the normal illumination image is used as the label image; Step A2: randomly crop each low-light image of size H×W×3 into an image of size P×P×3, and use the same random cropping method for its corresponding normal-light image to ensure that they have the same size and position, where H and W are the height and width of the low-light image and the normal-light image, and P is the height and width of the cropped image; Step A3: For each training paired image, randomly apply one of the following eight data augmentation methods: keep the original image, flip vertically, rotate 90 degrees, rotate 90 degrees and then flip vertically, rotate 180 degrees, rotate 180 degrees and then flip vertically, rotate 270 degrees, and rotate 270 degrees and then flip vertically.
3. The low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction according to claim 1, It is characterized in that The specific implementation steps of step B2 are as follows: Step B21: Input the feature X extracted by the mapping module in step B1 in , before being sent to the interactive enhancement block, two copies are first copied as the input of the local perception enhancement module and the global perception enhancement module respectively; here the input of the local perception enhancement module in the lth interactive enhancement block is represented as The input of the global perception enhancement module is expressed as Step B22, design a local perception enhancement module, which is mainly composed of channel context modeling, channel transformation, spatial context modeling and spatial transformation units; for the input feature D l-1 , the channel correction factor Z is obtained by channel context modeling and channel transformation l ; The spatial correction factor S is obtained by spatial context modeling and spatial transformation l ; Then, the channel correction factor Z l and the spatial correction factor S l Acting on the input feature D l-1 , to obtain the output feature D of the local perception enhancement module l ; The process is expressed by the formula: in, represents bitwise multiplication, Indicates bitwise addition; Step B23: Design local to global interactive operations for the output D of the local perception enhancement module. l , and input it into the global perception enhancement module as the search domain feature; Step B24: Design a global perception enhancement module, which mainly consists of a spatial cross-attention layer, a normalization layer, a channel self-attention layer, and a feed-forward network. For the input feature G l-1 , take it as the query domain feature, together with the search domain feature D l in step B23, and first send them into the spatial cross-attention and normalization layers in sequence. The obtained feature is then residual-connected to G l-1 to output an intermediate feature Then, send into the channel self-attention layer and the normalization layer in sequence. The obtained feature is then residual-connected to to output an intermediate feature Finally, send into the feed-forward network to obtain the output feature G l of the global perception enhancement module. This process is expressed by the formula: Among them, LN(·) represents the layer normalization operation, SMCA represents spatial cross attention, CMSA represents channel self-attention, and FFN represents feedforward network; Step B25, design a global to local interactive fusion module, which reuses the global perception enhancement module structure; for the output feature D of the local perception enhancement module l And the output module G of the global perception enhancement module l , respectively, and use them as the query domain and search domain, and send them to the spatial cross attention and normalization layers in turn. The obtained features are then combined with D l Residual connection, output intermediate features Then, The features are sent to the channel self-attention layer and normalization layer in turn, and then Residual connection, output intermediate features Finally, Send it to the feedforward network to get the final output feature X l , which is the output feature of the lth interaction enhancement block; the process is expressed by the formula:
4. The low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction according to claim 3, It is characterized in that The specific implementation steps of step B22 are as follows: Step B221: For input features First, it is processed by the channel context modeling unit; specifically, the features are first compressed by global average pooling to obtain the channel correction factor Next, the correction factor is concatenated with the historical l-1 channel correction factors to obtain the total correction factor representation It can be expressed as: Among them, AvgPool(·) represents global average pooling, [·;·] represents the concatenation operation; Step B222: Obtain the total correction factor for step B221 The channel transformation unit continues processing; it passes through a 1×1 convolution layer, a ReLu activation function, a 1×1 convolution layer, and a Sigmoid function in turn to obtain the current channel correction factor The process is expressed as: Where Conv1(·) represents a 1×1 convolutional layer, ReLu(·) represents a ReLu activation function, and Sigmoid(·) represents a Sigmoid function. Step B223: For input features It is also processed by the spatial context modeling unit; specifically, for D l-1 , firstly, the global average pooling and global maximum pooling operations along the channel dimension are performed to obtain a size of The global average feature map and the global maximum feature map are obtained, and the two feature maps are concatenated along the channel dimension; the process is expressed as follows: Among them, Avgpool(·) and MaxPool(·) represent global average pooling and global maximum pooling operations respectively. Represents the feature map after splicing; Step B224: feature map of step B223 The spatial transformation unit continues to process the image, which passes through a 1×1 group convolution layer, a 3×3 convolution layer, and the Tanh function to obtain the spatial correction factor. The formula is: Among them, GroupConv(·) represents a 1×1 group convolution layer, conv3(·) represents a 3×3 convolution layer, and Tanh(·) represents the Tanh function; Step B225: For input features First, compare it with the channel correction factor obtained in step B222 Multiply, and then add the spatial correction factor obtained in step B224 Add together to get the output D of the local perception enhancement module l .
5. The low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction according to claim 3, It is characterized in that The specific implementation steps of step B24 are as follows: Step B241, design a spatial cross attention layer, for the input feature G l-1 and D l , first reshaped into features of size C×N, where Then three groups of 1×1 convolutions are used to transform G l-1 Mapped to the "query" matrix Q sg , D l Mapped to the "key" matrix K sd and the "value" matrix V sd ; Further, the three matrices are each divided into h heads, each head has a size of N and a channel number So we get the "query" matrix for the i-th head "Key" Matrix and the "value" matrix where i∈[1,h]; Step B242: K sd,i Calculate softmax along the spatial dimension to get the weight of the spatial dimension Q sg,i Calculate softmax along the channel dimension to get the weight of the channel dimension Then CK sd,i Perform transpose operation and V sd,i Perform matrix multiplication to obtain the global information of N×N space, and then add it to SQ sg,i Multiply by the aggregated channel weight; finally, concatenate the features aggregated by all attention heads and perform a linear transformation with a 1×1 convolution to obtain the features enhanced by the spatial cross attention layer The process is expressed as: CK sd,i =softmax1(K sd,i ) SQ sg,i =softmax2(Q sg,i ) Among them, softmax1(·) represents the softmax function calculated along the spatial dimension, softmax2(·) represents the softmax function calculated along the channel dimension, [·;·] represents the concatenation operation, and W o Represents the learnable parameter matrix of 1×1 convolution; Step B243: Enhanced features of the spatial cross attention layer in step B242 After transposing it, it is sent to the normalization layer, and the obtained features are then combined with G l-1 Residual connection, the obtained features As the input of the next stage; expressed by the formula: Among them, Transpose(·) means transposing the C×N size feature to the N×C feature; Step B244, design the channel self-attention layer, for the input features Use three sets of 1×1 convolutions to map it into a "query" matrix Q cg , "key" matrix K cg and the "value" matrix V cg ; Split each of the three matrices into h heads, each head has a size of N and a number of channels So we get the "query" matrix for the i-th head "Key" Matrix and the "value" matrix where i∈[1,h]; Step B245: K cg,i Calculate softmax along the channel dimension to get the weight of the channel dimension Q cg,i Calculate softmax along the spatial dimension to get the weight of the spatial dimension Then CK cg,i Perform transpose operation and V sd,i Perform matrix multiplication to get a matrix of size Channel global information, then with SQ cg,i Multiply by the aggregated spatial weight; finally, concatenate the features aggregated by all attention heads and perform a linear transformation with a 1×1 convolution to obtain the features enhanced by the channel self-attention layer The process is expressed in the formula: CK cg,i =softmax2(K cg,i ) SQ cg,i =softmax1(Q sg,i ) Among them, softmax2(·) represents the softmax function calculated along the channel dimension, softmax1(·) represents the softmax function calculated along the spatial dimension, [·;·] represents the concatenation operation, and W p Represents the learnable parameter matrix of 1×1 convolution; Step B246: Enhanced features of the channel cross attention layer in step B245 Send it to the normalization layer, and then get the features Residual connection to obtain features Step B247: Send it to the feedforward network to get the output feature G of the global perception enhancement module l .
6. The low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction according to claim 5, It is characterized in that The specific implementation steps of step B247 are as follows: Step B2471: Features enhanced by channel self-attention First, the channel is doubled by a 1×1 convolution, and then a 3×3 depth-separable convolution is used for feature mapping; then, the mapped features are divided into two parts along the channel dimension, represented as and The process is expressed as: Among them, Conv1(·) represents a 1×1 convolutional layer, DWConv3(·) represents a 3×3 depth-separable convolutional layer, and chunk(·) means dividing the features into two parts along the channel dimension; Step B2472: and The two features are activated by ReLu and multiplied by each other’s original features to achieve information interaction. Then, a 1×1 convolution is used to transform the features to obtain the output feature G of the global perception enhancement module. l ; The process can be expressed as: Where ReLu(·) represents the ReLu activation function, W t Represents the learnable parameters of the 1×1 convolution.
7. The low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction according to claim 1, It is characterized in that The specific implementation of step C is: Step C: Design the loss function, using L1 loss Structural loss function l structure and VGG perceptual loss l perceptual Composition, the total objective loss function of the network l total It is expressed as follows: l perceptual =||(VGG 3,8,15 (I en )-VGG 3,8,15 (I gt ))|| 2 Among them, λ 1 , 2 and λ 3 is the equilibrium parameter, I en For enhanced images, I gt is the image with normal illumination; μ en and μ gt is the mean of the two images, σ en and σ gt represents the variance of the two images, c 1 and c 2 are two constant values to prevent the denominator from being 0; ||·|| 2 Indicates the calculation of mean square error, VGG 3,8,15 (·) indicates that the features of layer 3, layer 8, and layer 15 are extracted using the VGG-16 classification model pre-trained on the ImageNet dataset.
8. The low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction according to claim 1, It is characterized in that The specific implementation steps of step D are as follows: Step D1, randomly divide the training data set obtained in step A into several batches, each batch containing N pairs of images; Step D2: input low-light image I in After the interactive enhancement network in step B, the enhanced image I is obtained en , use the formula in step C to calculate the loss l total ; Step D3: Calculate the gradient of the parameters in the network using the back propagation method according to the loss, and update the network parameters using the Adam optimization method; Step D4: Repeat steps D1 to D3 in batches to obtain an interactive enhancement network model.
9. The low-light image enhancement method based on CNN-Transformer hybrid modeling and bilateral interaction according to claim 1, It is characterized in that The specific implementation steps of step E are as follows: Step E1: Input a low-light image I of size H×W×3 in Reshape the image into a size of P×P×3, where H and W are the height and width of the low-light image, respectively, and P is the height and width of the reshaped image; Step E2, sending the reshaped image to the interactive enhancement network in step B to obtain an enhanced image of size P×P×3; Step E3: Reshape the enhanced image of size P×P×3 into an enhanced image I of size H×W×3 en .