Zero-reference low-illumination image enhancement method based on residual quotient learning

By adopting a residual quotient learning method in low-illumination image enhancement, using the enhancement network and calibration network to combine residual blocks and attention mechanisms, the problem of poor image enhancement effect in the prior art is solved, and a higher quality low-illumination image enhancement is achieved.

CN120070214APending Publication Date: 2025-05-30NANJING FORESTRY UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510103487.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-15
Filing Date
2025-01-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate the channel information and spatial information of the image in low-illumination image enhancement, resulting in insufficient images after the enhancement in some aspects.

Method used

Using a zero-reference low-illumination image enhancement method based on residual quotient learning, the image features are extracted and weighted by building an enhancement network and calibration network, using residual blocks and advanced attention mechanisms such as RC-CBAM, and additional feature information is provided through the calibration network to accelerate training.

Benefits of technology

It achieves a more natural and visual effect low-illumination image enhancement, improves the model's ability to process complex lighting changes, and the generated enhanced images perform excellently in brightness, contrast, detail clarity and semantic information retention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070214A_ABST
    Figure CN120070214A_ABST
Patent Text Reader

Abstract

The invention discloses a low-illumination image enhancement method. Firstly, a low-illumination image is acquired and preprocessed into a training set, then a neural network constructed based on a residual quotient learning framework is trained by using the training set, and the network comprises an enhancement network and a calibration network. And processing the input image by using the enhanced network after training. According to the method, feature information can be deeply mined, and channel and space information can be integrated. The core of the method is a residual quotient learning framework, an enhancement task is converted into prediction of a potential quotient through residual learning, and self-adaptive adjustment is carried out according to an original low-illumination image. The frame improves the capability of processing complex light, enables the generated image to be more natural and better in visual effect, achieves efficient and high-quality enhancement, and has wide application prospects and practical values in a plurality of fields. A low-illumination enhancement neural network based on residual quotient learning is constructed, and a learning framework converts a low-illumination enhancement task into prediction of a potential quotient (ratio) adaptively according to an original low-illumination input image through residual learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a zero-reference low-light image enhancement method based on residual quotient learning. Specifically, a specific neural network framework is used to perform low-light enhancement operations on images to improve the visual quality and usability of the images, making them more suitable for computer vision tasks, image display and other application scenarios with high requirements on image quality. Background Art

[0002] In the wide application scenarios of modern image acquisition and processing technology, the low-light image problem has always been a key challenge that needs to be solved. Images acquired in low-light environments often have many defects, which seriously affect the subsequent application value of the images.

[0003] With the rapid development of deep learning technology, it has shown great potential and advantages in the field of image processing. Deep learning models can automatically learn complex features and patterns in images through large-scale data training, thereby achieving more accurate and adaptive image enhancement. In the field of low-light image enhancement, there have been some research works based on deep learning. For example, some studies use simple convolutional neural networks to directly process low-light images. Although certain effects have been achieved, due to the relatively simple network structure, the extraction and utilization of image features are not sufficient. When processing complex low-light images, it is still difficult to achieve satisfactory enhancement effects. Some other studies have tried to introduce attention mechanisms into the network to increase the network's attention to key areas and features of the image, but these methods are often not perfect in the design of the attention mechanism, and cannot fully and effectively integrate the channel information and spatial information of the image, resulting in the enhanced image still having deficiencies in some aspects. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a zero-reference low-light image enhancement method based on residual quotient learning in view of the deficiencies of the above-mentioned prior art. The zero-reference low-light image enhancement method based on residual quotient learning fully utilizes the deep feature extraction capability and image detail enhancement capability of the enhancement network by constructing an enhancement network and a calibration network. Among them, the attention module can automatically learn and assign different weights to features of different channels and spatial positions, so as to more comprehensively and effectively integrate the channel information and spatial information of the image. The residual quotient is obtained by combining the low-light image and the output of the enhancement network, so as to better adapt to the non-uniform lighting conditions commonly seen in natural low-light images. It effectively improves the model's ability to handle complex lighting changes, thereby generating more natural and visually enhanced images; at the same time, combined with the feature assistance and accelerated training of the calibration network, the network can deeply mine various feature information in low-light images, which is helpful to extract rich multi-level features. The calibration network can extract additional image feature information during the training process through its unique input convolutional layer, multiple intermediate convolutional layers and output convolutional layer structure, overcoming the limitations of traditional low-light image enhancement methods, achieving more efficient and high-quality low-light image enhancement effects, and providing strong technical support and guarantee for many fields that rely on high-quality image data.

[0005] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:

[0006] A zero-reference low-light image enhancement method based on residual quotient learning comprises the following steps:

[0007] Step 1: Obtain low-light images and preprocess them as training sets;

[0008] Step 2: construct a low-light enhancement neural network based on a residual quotient learning framework, which includes an enhancement network and a calibration network; during image training, the enhancement network, the residual quotient and the calibration network are simultaneously used to enhance the input image;

[0009] Step 3: Train the constructed low-light enhancement neural network based on the residual quotient learning framework through the training set to obtain a trained low-light enhancement neural network model;

[0010] Step 4: Use the enhanced network in the trained low-light enhancement neural network model to process the input low-light image to obtain the final image.

[0011] As a further improved technical solution of the present invention, the method for low-light image enhancement based on the residual quotient learning framework mainly consists of an enhancement network and a calibration network, which work together to achieve low-light image enhancement. During the training process, the enhancement network first processes the input image, extracts additional feature information, and then passes this information to the calibration network.

[0012] The enhanced network includes a neural network model improved based on the residual network framework, which introduces residual blocks (i.e., SE-RESB modules) and combines components such as the convolutional block attention mechanism (RC-CBAM module). Each residual block (i.e., SE-RESB module) includes four consecutive convolutional layers, each followed by a batch normalization layer and a ReLU activation function. One SE-RESB module is used to weight the channels of the feature map. The RC-CBAM attention module includes a channel attention module and a spatial attention module, and uses multiple residual connection operations on the channel attention module and the spatial attention module to combine the weights of the channel and spatial positions. The enhanced network also includes a Laplacian sharpening module.

[0013] The calibration network includes an input convolutional layer, multiple intermediate convolutional layers, and an output convolutional layer. Among them, the input convolutional layer, as the first layer of the calibration network, performs convolutional operations using Conv2d, performs batch normalization processing with BatchNorm2d, and introduces non-linearity with the ReLU activation function. Two intermediate convolutional layers: Each intermediate convolutional layer contains a convolutional layer, a batch normalization layer, and a ReLU activation function inside. The convolutional layer adopts the structure of Conv2d and cooperates with the batch normalization layer and the activation function to form a complete feature extraction and transformation unit. The output convolutional layer: Located at the last layer of the calibration network, it converts the feature map processed by the previous multiple convolutional layers into an output of a specific dimension.

[0014] The neural network module ResNet improved based on the residual network framework in the enhancement network includes an initial convolutional layer sequence, a residual block sequence, and a final convolutional layer sequence. Initial convolutional layer sequence: It contains a two-dimensional convolutional layer with an input channel number of 3 (corresponding to the RGB channels of the input image), an output channel number of 3, which is used for preliminary feature extraction and channel conversion of the input image; a batch normalization layer that can accelerate network convergence and improve the generalization ability of the model; a ReLU activation function that introduces non-linearity and enhances the expression ability of the network; a convolutional block attention module RC-CBAM with an input channel number of 3; and finally a max pooling layer with a convolutional kernel size of 3. Residual block sequence: It consists of multiple residual blocks (i.e., SE-RESB modules). Each residual block (i.e., SE-RESB module) includes four consecutive convolutional layers with a convolutional kernel size of 3, a stride of 1, and a padding of 1. After each convolutional layer, there is a batch normalization layer and a ReLU activation function. A SE layer (channel attention module) is also included in the residual block. It first compresses the feature map in the spatial dimension through adaptive average pooling to obtain a global feature description in the channel dimension, and then passes through a network structure composed of two fully connected layers, with ReLU activation in the middle and finally obtaining the weight of each channel through the Sigmoid function, which is used to weight the channels of the feature map. The design of the residual block adopts the idea of residual connection, that is, the input feature map is saved as the part of the residual connection. After a series of convolutional, normalization, and activation operations, the obtained feature map is added to the residual connection part to obtain the final output feature map of the residual block. Final convolutional layer sequence: It includes a two-dimensional convolutional layer with an input channel number of 3, an output channel number of 3, a convolutional kernel size of 3, a stride of 1, and a padding of 1, which is used to further adjust the channel information of the feature map; a batch normalization layer and a ReLU activation function to maintain the non-linear characteristics of the network.

[0015] As a further improved technical solution of the present invention, in step 2:

[0016] The processing process of the enhancement network is as follows:

[0017] illu = Enhance(input);

[0018] Among them, Enhance represents the enhancement network, input is the input image of the enhancement network, and illu is the output image of the enhancement network;

[0019] The processing process of the residual quotient is as follows:

[0020] Calculate the brightness adjustment factor r through the output image illu of the enhancement network:

[0021]

[0022] torch.clamp(r, 0, 1) means to limit the brightness adjustment factor r between 0 and 1 to obtain the brightness adjustment factor r; the brightness adjustment factor r is the residual quotient;

[0023] The processing process of the calibration network is as follows:

[0024] Send the brightness adjustment factor r into the calibration network;

[0025] att = Calibrate(r);

[0026] In the formula, Calibrate represents the calibration network, and att is the output image of the calibration network.

[0027] As a further improved technical solution of the present invention, in step 2, the specific processing process of the enhancement network is as follows:

[0028] Step 2a: Process the input image data input through the ResNet module to obtain the feature map fea l , the formula is:

[0029] fea l = ResNet(input);

[0030] Where ResNet represents the improved residual neural network;

[0031] Step 2b: Sharpen the obtained feature map fea l through the Laplacian sharpening module, and the formula is:

[0032] fea = LaplacianSharpening(fea l );

[0033] In the formula, LaplacianSharpening represents the Laplacian sharpening module;

[0034] Step 2c: Calculate the enhanced image illu:

[0035] illu = fea * factor + input;

[0036] That is, the enhanced image illu is obtained by multiplying the sharpened feature map fea by the learnable parameter factor and then adding it to the original input image input;

[0037] Step 2d: Limit the numerical range of the calculated illu, limit its value between 0.0001 and 1, to obtain the final output image of the enhancement network.

[0038] As a further improved technical solution of the present invention, in the step 2a, the specific processing process of the improved Residual Neural Network ResNet module is as follows:

[0039] Step 2a1: If the input image data input is denoted as x, the input image data x is processed through the initial convolutional layer sequence initial_conv, and the formula is:

[0040] x 1 = initial_conv(x);

[0041] The initial convolutional layer sequence initial_conv includes a two-dimensional convolutional layer, a batch normalization layer, a ReLU activation function, a convolutional block attention module RC-CBAM, and a max pooling layer with a convolutional kernel size of 3×3;

[0042] Step 2a2: The feature map x 1 is processed through the residual block sequence SE-RESB module for three cycles to obtain the processed feature map a 3 , and the formula is:

[0043] a 3 = residualblocks(residualblocks(residualblocks(x 1 )));

[0044] In this formula, residualblocks is the SE-RESB module;

[0045] Step 2a3: The feature map a obtained through the residual block sequence SE-RESB module 3 is processed through the final convolutional layer sequence final_conv for two cycles, and the formula is:

[0046] x 3 = final_conv(final_conv(a 3 ));

[0047] In the formula, final_conv is the final convolutional layer sequence, and x 3 is the feature map obtained through two cycles of processing by the final convolutional layer sequence final_conv;

[0048] The final convolutional layer sequence final_conv includes a two-dimensional convolutional layer, a batch normalization layer, and a ReLU activation function.

[0049] As a further improved technical solution of the present invention, in the step 2a2, the specific processing process of the SE-RESB module is as follows:

[0050] Step 2a21: Save the input feature map x 1 as part of the residual connection;

[0051] Step 2a22: Perform the following operations on the input feature map x 1 :

[0052] a = ReLU(BatchNorm(Conv2d(x 1 )));

[0053] Conv2d represents a two-dimensional convolution operation, BatchNorm represents a batch normalization operation, ReLU represents a rectified linear unit activation function, and a is the new feature map obtained after processing;

[0054] Step 2a23: The feature map a goes through the operations in Step 2a22 and loops 3 more times to obtain the processed feature map a 1 ;

[0055] The feature map a 1 is further processed by the channel attention mechanism SE layer, and the formula is:

[0056] a 2 = SELayer(a 1 );

[0057] In the formula, SELayer represents the SE layer;

[0058] Add the feature map a processed by the channel attention mechanism SE layer 2 to the residual connection part x saved in Step 2a21 1 , and the formula is:

[0059] a 3 = a 2 + x 1 ;

[0060] Thus, the final output feature map a of the SE-RESB module is obtained 3 .

[0061] As a further improved technical solution of the present invention, in Step 2a1, the convolutional block attention module RC-CBAM (CBAM attention mechanism) mainly consists of a channel attention module and a spatial attention module.

[0062] The processing process of the convolutional block attention module RC-CBAM is as follows:

[0063] Step 2a11: Denote z as the input feature map of the convolutional block attention module RC-CBAM;

[0064] Process the input feature map z through the channel attention module Cam to obtain the feature map z after channel attention adjustment 1 , the formula is:

[0065] z 1 = Cam(z);

[0066] Step 2a12: Process the feature map z obtained in the previous step 1 through the spatial attention module Sam, and the formula is:

[0067] z 2 = Sam(z 1 );

[0068] Step 2a13: Process the feature map z again 2 through the channel attention module Cam and add it to the original feature map z 2 , and the formula is:

[0069] z 3 = Cam(z 2 ) + z 2 ;

[0070] Step 2a14: Process the feature map z again 3 through the spatial attention module Sam and add it to the original feature map z 3 , and the formula is:

[0071] z 4 = Sam(z 3 ) + z 3 ;

[0072] Step 2a15: Process the feature map z again 4 through the channel attention module Cam to obtain:

[0073] z 5 = Cam(z 4 );

[0074] Step 2a16: Finally, process the feature map z 5 through the spatial attention module Sam to obtain the final output feature map z 6 :

[0075] z 6 = Sam(z 5 ).

[0076] As a further improved technical solution of the present invention, in the step 2a1, for the channel attention module: first, an adaptive max pooling layer and an adaptive average pooling layer are created, and the input feature map is respectively pooled into a feature map of size 1×1 to obtain two different feature descriptions. Then, each passes through a network structure including a linear layer, a Tanh activation function, a linear layer, and a Sigmoid function.

[0077] The processing process of the channel attention module Cam is as follows:

[0078] (1). Create an adaptive max pooling layer AdaptiveMaxPool2d, which pools the input feature map Q into a feature map y of size 1×1 maxpool , and the formula is:

[0079] y maxpool = AdaptiveMaxPool2d(1)(Q);

[0080] In the formula, Q is the input feature map, and AdaptiveMaxPool2d(1)(Q) represents the adaptive max pooling operation of the input feature map Q to size 1;

[0081] Then, it passes through a sequence composed of multiple fully connected layers, specifically including:

[0082] (1a). The first fully connected layer Linear expands the number of input channels to obtain Q 1 , and the formula is:

[0083] Q 1 = Linear(inChannel, int(inChannel*r))(y maxpool );

[0084] In the formula, inChannel represents the number of input channels, y maxpool is the input after being flattened by the max pooling layer, Linear represents the fully connected layer operation, and int(inChannel*r) represents the operation of expanding the number of input channels by r times;

[0085] (1b). Then, it is processed by the hyperbolic tangent activation function Tanh, and the formula is:

[0086] Q 2 = Tanh(Q 1 );

[0087] (1c). The second fully connected layer Linear restores the dimension of the feature map Q 2 to the original input size, and the formula is:

[0088] Q 3= Linear(int(inChannel * r), inChannel)(Q 2 );

[0089] (1d), finally, after being processed by the Sigmoid activation function, the weight weight of the max pooling branch is obtained max , and the formula is: weight max = Sigmoid(Q 3 );

[0090] (2). Create an adaptive average pooling layer AdaptiveAvgPool2d, and the formula is:

[0091] y avgpool = AdaptiveAvgPool2d(1)(Q);

[0092] AdaptiveAvgPool2d(1)(Q) represents an adaptive average pooling operation that reduces the input feature map Q to a size of 1;

[0093] Then, it passes through a sequence composed of multiple fully connected layers, specifically including:

[0094] (2a). The first fully connected layer Linear expands the number of input channels, and the formula is:

[0095] q 1 = Linear(inChannel, int(inChannel * r))(y avhpool );

[0096] In the formula, inChannel represents the number of input channels, y avgpoil is the input after being flattened by the average pooling layer, Linear represents the fully connected layer operation, and int(inChannel * r) represents the operation of expanding the number of input channels by r times;

[0097] (2b). Then, it is processed by the hyperbolic tangent activation function Tanh, and the formula is:

[0098] q 2 = Tanh(q 1 );

[0099] (2c). The second fully connected layer Linear restores the dimension of the feature map q 2 to the original input size, and the formula is:

[0100] q 3 = Linear(int(inChannel * r), inChannel)(q 2 );

[0101] (2d), finally, after being processed by the Sigmoid activation function, the weight of the average pooling branch, weight, is obtained avg , and the formula is: weight avg = Sigmoid(q 3 );

[0102] (3) Add the weight of the max pooling branch, weight max and the weight of the average pooling branch, weight avg to obtain the fused weight, weight. The formula is:

[0103] weight = weight max + weight avg ;

[0104] (4) Reshape the fused weight, weight, into a shape matching the input feature map Q. Assuming the shape of the weight tensor is (h, w), the reshape formula is:

[0105] Mc = reshape(weight, (h, w, 1, 1));

[0106] In the formula, reshape is the reshape operation, Mc is the reshaped weight, and the tensor of weight is reshaped from a two-dimensional tensor to a new four-dimensional tensor (h, w, 1, 1), where: h and w represent the height and width of the new tensor respectively; 1 and 1 indicate that the last two dimensions of the new tensor are single-element dimensions;

[0107] Finally, multiply the reshaped weight, Mc, element-wise with the input feature map Q to obtain the feature map z after channel attention adjustment i , and the formula is:

[0108] z i = Mc * Q.

[0109] As a further improved technical solution of the present invention, in the step 2a1, the processing process of the spatial attention module Sam is as follows:

[0110] Take the maximum value and the average value of the input feature map l in the channel dimension respectively, and then concatenate these two results in the channel dimension to obtain a new feature map m. The formula is:

[0111] m = torch.cat((max(l), mean(l));

[0112] In this formula, torch.cat is the concatenation operation, max(l) is the operation of taking the maximum value of the feature map l, and mean(l) is the operation of taking the average value of the feature map l;

[0113] Next, the new feature map m passes through a convolutional layer with a convolutional kernel size of 7 and a padding of 3 and the Sigmoid function to obtain the weights at each spatial position, which is used to reflect the importance of features at different spatial positions.

[0114] Finally, the obtained spatial weights Ms are multiplied element-wise with the input feature map l to obtain the feature map t after spatial attention adjustment. The formula is:

[0115] t = Ms * l.

[0116] As a further improved technical solution of the present invention, the processing process of the calibration network is as follows:

[0117] Denote the input as the brightness adjustment factor r and the output as the image att;

[0118] First, the brightness adjustment factor r is input into the input convolutional layer for feature extraction to obtain the feature map F 1 , and the formula is:

[0119] F 1 = ReLU(BatchNorm(Conv2d(r)));

[0120] Conv2d represents a two-dimensional convolution operation, BatchNorm represents a batch normalization operation, and ReLU represents a rectified linear unit activation function;

[0121] The feature map F 1 is input into the intermediate convolutional layer for feature extraction to obtain the feature map F 2 , and the feature map F 2 is input into the second intermediate convolutional layer for feature extraction to obtain the feature map F 3 , and the formula is:

[0122] F 2 = ReLU(BatchNorm(Conv2d(F 1 )));

[0123] F 3 = EeLU(BatchNorm(Conv2d(F 2 )));

[0124] Then, the brightness adjustment factor r is subtracted from the feature map F 3 to obtain the feature map The formula is:

[0125]

[0126] The feature map goes through the output convolutional layer operation to obtain the feature map The formula is:

[0127]

[0128] Perform a residual connection addition operation on the feature map and the brightness adjustment factor r. The formula is:

[0129]

[0130] The network framework trained by the low - illumination image enhancement method is a zero - reference low - illumination image network for residual quotient learning. The specific steps are as follows:

[0131] First, initialize the weights of the convolutional layer. The formula is:

[0132] weight.data ∼ N(0, 0.02 2 );

[0133] In the formula, weight.data represents the weights, and N(μ, σ 2 ) represents a normal distribution with mean μ and standard deviation σ, that is, the weights are initialized to a normal distribution with mean 0 and standard deviation 0.02. Initialize the weights of the batch normalization layer.

[0134] For the batch normalization layer of the BatchNorm2d type, the weight initialization formula is:

[0135] weight.data ∼ N(1, 0.02 2 );

[0136] That is, the weights are initialized to a normal distribution with mean 1 and standard deviation 0.02.

[0137] After the weight initialization is completed, process the current input image through the enhancement network:

[0138] illu = Enhance(input);

[0139] In the formula, Enhance represents the enhancement network, which includes a residual network, an attention mechanism, and a Laplacian sharpening operation to perform feature extraction, enhancement, and sharpening on the input image, and then outputs the enhanced image illu.

[0140] Next, further calculate the brightness adjustment factor r:

[0141]

[0142] r = torch.clamp(r, 0, 1);

[0143] In the formula, the brightness adjustment factor r (i.e., the image r after brightness adjustment) is obtained by dividing the input feature map input by illu. torch.clamp is used to limit the brightness adjustment factor between 0 and 1 to obtain the image r after brightness adjustment. This step aims to adjust the brightness of the original input image according to the enhanced image to make the image clearer and more natural.

[0144] The image r after brightness adjustment is fed into the calibration network:

[0145] att = Calibrate(r);

[0146] In the formula, Calibrate represents the calibration network, which enhances the image according to its internal structure (such as multi-layer convolution, batch normalization, and residual connection, etc.).

[0147] During the training process, the results of the enhanced images at each stage need to be recorded simultaneously. Initialize the loss value loss to 0. Then iterate through each loop stage, and calculate the loss value loss between the input image in[i + 1] and the enhanced image in[i] at each loop stage through the loss function object Criterion i , loss i represents the loss value corresponding to the i-th loop stage, and the loss values loss i of multiple loop stages are accumulated into the total loss loss. Through the obtained total loss loss, then the image passes through the training framework again until the target total loss value is reached to complete the training.

[0148] During the testing process, the initialization of the convolutional layer weights and the initialization of the batch normalization layer weights are the same as those in the training process. If the input image is input p , first, it is processed sequentially through the ResNet module and the Laplacian sharpening module in the enhancement network to obtain the feature representation p. The formula is: p = LaplacianSharpening(ResNet(input p )); Then, the enhanced feature p is added to the original input image input p : p 1 = p + input p ; In the formula, p 1 is the feature map output after combination.

[0149] Then, the brightness adjustment factor p 2 is obtained in the same way as in the training process. The formula is:

[0150]

[0151] Finally, the model returns the image p after brightness adjustment 2, this result is the enhanced image during the test process (application process), which can be used for subsequent analysis, evaluation, or further processing.

[0152] The loss function used for training the zero-reference low-light image enhancement method based on residual quotient learning in the present invention is composed of the addition of a fidelity loss, a smoothness loss, a color loss, and a perceptual loss. The formula is:

[0153] Total loss =1.5×Fidelity loss +Smooth loss +Color loss +Perception loss ;

[0154] Where Fidelity loss represents the fidelity loss, which is the square root of the square of the previous image minus the currently generated image. Smooth loss represents the smoothness loss, Color loss represents the color loss, and Perception loss represents the perceptual loss.

[0155] 1) The structure of the color loss is:

[0156] Assume the input image is x, with a shape of (b, c, h, w), where b represents the batch size, c represents the number of channels (3 for color images, i.e., RGB channels), and h and w represent the height and width of the image respectively.

[0157] First, calculate the average RGB value mean_rgb of the image in the spatial dimensions (height and width) b,c,0,0 , and the formula is:

[0158]

[0159] The mean_rgb obtained from the formula b,c,0,0 has a shape of (b, 3, 1, 1). Then, through the torch.split operation, split the average RGB value into three separate tensors mr b (red channel mean), mg b (green channel mean), and mb b (blue channel mean) in the channel dimension, and the shape of each tensor is (b, 1, 1, 1).

[0160] For the i-th image in the batch, its average RGB value is calculated as follows:

[0161]

[0162] In the formula, mrb Represents the average value of the red channel, mg b Represents the average value of the green channel, mb b Represents the average value of the blue channel, and then calculate the squared difference between the average values of different color channels:

[0163] Drg = (mr b - mg b ) 2 ;

[0164] In the formula, Drg is used to measure the difference between the average values of the red and green channels.

[0165] Drb = (mr b - mb b ) 2 ;

[0166] In the formula, Drb is used to measure the difference between the average values of the red and blue channels.

[0167] Dgb = (mb b - mg b ) 2 ;

[0168] In the formula, Dgb is used to measure the difference between the average values of the blue and green channels.

[0169] To improve numerical stability, add a small constant epsilon = 1e-8 and calculate the color loss k:

[0170]

[0171] This color loss aims to make the color distribution of the enhanced image more uniform and natural, avoiding excessive color deviation.

[0172] 2) Smooth loss:

[0173] Denote the input image as input and output, representing the original low-light image and the enhanced image (the result of the training framework, i.e., the result after passing through the enhancement network and the calibration network). In this module, mainly input is used for calculation, but output is used for operations related to the gradient calculation of input.

[0174] First, convert the input image input from the RGB color space to the YCbCr color space using the rgb2yCbCr method. In this method, first flatten the input image into a two-dimensional tensor with a shape of (-1, 3), then multiply it with a predefined conversion matrix and add a bias term, and finally reshape the result to the same shape as the input image.

[0175] Then calculate the pixel gradient weights of the luminance channel (Y) in the YCbCr color space. For the differences between adjacent pixels in the horizontal and vertical directions and adjacent pixels on the diagonal, calculate the weights w1 to w24 respectively. Taking w1 as an example, the calculation formula is:

[0176]

[0177] where σ is the parameter defined during initialization, (input[:,c,1:,:]-input[:,c,:-1,:]) 2 calculates the difference between adjacent pixels in the luminance channel, and converts the difference into a weight through an exponential function. The greater the difference, the smaller the weight. The calculations of the other weights w2 to w24 are similar, corresponding to the pixel differences in different directions and distances respectively.

[0178] According to the calculated weights, calculate the pixel gradient penalty terms of the output image output in different directions. Taking the pixel gradient penalty term pixel g rad1 in the first direction as an example, the calculation formula is:

[0179]

[0180] where p = 1.0 indicates using the L1 norm to calculate the pixel gradient, w1 is the corresponding weight, and |output[:,c,1:,:]-output[:,c,:-1,:]| p calculates the difference between adjacent pixels in the output image, and obtains the pixel gradient penalty term pixel g rad1 in different directions by multiplying by the weight w1.

[0181] 3) Perceptual loss:

[0182] The perceptual loss module first obtains the feature extraction part of the pre-trained VGG16 network. Then, four sequential modules are created: l 1 、l 2 、l 3 、l 4 .

[0183] where l 1 adds the features of the first 4 layers of the VGG16 network to this module in sequence, implemented by looping through layers 1 - 4; l 2 adds the features of layers 4 to 9 of the VGG16 network to this module, with the loop range being layers 4 - 9; l 3 adds the features of layers 9 to 16 of the VGG16 network to this module, with the loop range being layers 9 - 16; l 4The features of the 16th to 23rd layers of the VGG16 network are added to this module, with the loop range being layers 16 - 23. Finally, by traversing all the parameters of the module, the attributes of the parameters are set to False, freezing the parameters of these pre-trained layers, that is, not calculating their gradients and only using the features they extract to calculate the perceptual loss. The input image is x, which is first sent to the l 1 module for processing. This module contains the first 4 layers of the VGG16 network. Let the processing functions of these 4 layers be f 0 、f 1 、f 2 、f 3 respectively. Then the feature representation after passing through this module can be calculated by the formula: h 1 = f 3 (f 2 (f 1 (f 0 (x)))));

[0184] This step performs preliminary feature extraction on the input image, extracting low-level features of the image, such as edge and texture information. Then, h 1 is sent to the l 2 module, which contains the 4th to 9th layers of the VGG16 network. Let the processing functions of these 6 layers be g 0 、g 1 、g 2 、g 3 、g 4 、g 5 respectively. Then the feature representation h 1 is updated to h 2 : h 2 = g 5 (g 4 (g 3 (g 2 (g 1 (g 0 (h 1 ))))));

[0185] This step further extracts the intermediate features of the image, containing more semantic information. Then, h 2 is sent to the l 3 module, which contains the 9th to 16th layers of the VGG16 network. Let the processing functions of these 8 layers be p 0 、p 1 、p 2 、p 3 、p 4 、p 5 、p 6 、p 7 respectively. Then it is updated to h 2 : h 3= p 7 (p 6 (p 5 (p 4 (p 3 (p 2 (p 1 (p 0 (h 2 ))))))));

[0186] This step continues to extract higher-level semantic features for a deeper understanding of the overall structure and content of the image. Finally, h 3 is fed into the l 4 module, which contains the 16th to 23rd layers of the VGG16 network. Let the processing functions of these 8 layers be q 0 , q 1 , q 2 , q 3 , q 4 , q 5 , q 6 , q 7 , then the finally obtained feature representation h 4 is: h 4 = q 7 (q 6 (q 5 (q 4 (q 3 (q 2 (q 1 (q 0 (h 3 ))))))));

[0187] This feature representation contains the high-level semantic information of the image and can reflect the overall perceptual features of the image. The model returns h 4 , and this feature representation will be used for the calculation of the subsequent perceptual loss.

[0188] Finally, the total loss is calculated, and the loss function is composed of the addition of the fidelity loss, smooth loss, color loss, and perceptual loss.

[0189] The formula is:

[0190] Total loss = 1.5 × Fidelity loss + Smooth loss + Color loss + Perception loss .

[0191] Another object of the present invention is to provide a computer system comprising a memory and a processor, wherein the memory is used to store computer programs / instructions; and the processor is used to execute the computer programs / instructions to implement the zero-reference low-illumination image enhancement method based on residual quotient learning.

[0192] The beneficial effects of the present invention are:

[0193] This zero-reference low-light image enhancement method based on residual quotient learning builds an enhancement network containing residual blocks and advanced attention mechanisms (such as channel attention and spatial attention modules and SE layers in RC-CBAM), enabling the network to deeply mine various feature information in low-light images. The combination of multiple convolutional layers, batch normalization layers and activation functions in the residual block helps to extract rich multi-level features, while the attention mechanism can automatically learn and give different weights to features of different channels and spatial positions, which can achieve comprehensive and effective integration of channel information and spatial information of the image. A residual learning framework is designed, that is, the residual quotient information is obtained by dividing the predicted low-light image with the output of the enhancement network. This measure can better adapt to the non-uniform lighting conditions commonly seen in natural low-light images, effectively improve the model's ability to handle complex lighting changes, and thus generate more natural and visually enhanced images. A calibration network is designed to work in conjunction with the enhancement network. The calibration network can extract additional image feature information during the training process through its unique input convolutional layer, multiple intermediate convolutional layers and output convolutional layer structure. In the test phase, a simple element-by-element division operation is used to obtain the residual quotient information. The final enhanced image can be quickly generated, which significantly speeds up the reasoning process and meets the efficiency requirements in practical applications.

[0194] The residual quotient of the present invention plays a key role in image enhancement. It is an important bridge connecting the low-light input image and the desired restored image, which is specifically manifested in the following aspects: Redefining the enhancement task: The residual quotient x is the relationship between the low-light image input and the enhanced network output image k. Reconstructed to transform the enhancement problem into the learning and prediction of the residual quotient. This redefinition provides a new perspective and solution idea for image enhancement, no longer limited to the traditional decomposition-enhancement-reconstruction process. Adapt to non-uniform illumination: Different from the common practice in existing Retinex-related methods of reducing the number of channels of the quotient from three (color) to one (grayscale), this method restores the three-channel structure. This measure can better adapt to the non-uniform illumination commonly found in natural low-light images, effectively improving the model's ability to handle complex lighting changes, thereby generating more natural and better visually enhanced images. Facilitate network learning and optimization: When designing the network, a neural network is used to learn the mapping from low-light images to the residual quotient. Since and are visually somewhat similar, assuming there is a relatively simple and easy-to-optimize connection between them, by transforming the direct mapping into learning the residual function (i.e., ), the network training process can be further simplified and the performance of the overall model can be improved. In the test stage, the final enhanced image can be quickly generated through a simple element-wise division operation , significantly accelerating the inference speed and meeting the efficiency requirements in practical applications.

[0195] The present invention constructs a neural network module improved based on the residual network framework, and introduces residual blocks and advanced convolutional block attention mechanisms (such as the channel attention and spatial attention modules and the SE layer in RC-CBAM) into the enhancement network. The combination of multiple convolutional layers, batch normalization layers, and ReLU activation functions in the residual blocks can effectively extract rich multi-level features, enabling the network to deeply learn the detail and structure information of the image. Among them, the neural network module improved based on the residual network framework in the enhancement network effectively extracts and accurately weights image features through components such as residual blocks and convolutional block attention mechanisms (RC-CBAM); the Laplacian sharpening module is used to enhance image details.

[0196] The calibration network consists of an input convolutional layer, multiple convolutional blocks, and an output convolutional layer, providing additional feature assistance for the enhancement network and accelerating training. A comprehensive loss function including color loss, perceptual loss, smoothness loss, and fidelity loss is also proposed. Experimental results show that this technology can effectively improve the brightness, contrast, detail clarity, and semantic information retention of low-light images, achieving a good balance between the number of parameters and performance, being applicable to multiple fields, and having broad application prospects and significant practical value.

[0197] The zero-reference network based on residual quotient learning designed by the present invention achieves a good balance between network parameters and performance. After experimental verification, the model performs well in improving image quality, and can effectively enhance the brightness, contrast, detail clarity of low-light images and retain the semantic information of images. Compared with traditional methods and some existing deep learning methods, under the same or similar parameter amounts, the present invention can obtain better image enhancement effects, such as higher peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) indicators, indicating that the generated enhanced images are closer to real high-quality images in visual quality, and have stronger adaptability and stability in different types of low-light images and actual application scenarios.

[0198] This zero-reference low-light image enhancement method based on residual quotient learning builds an enhancement network and a calibration network, making full use of the deep feature extraction capability of the residual network, the precise feature weighting capability of the convolutional block attention mechanism, and the image detail enhancement capability of the Laplace sharpening module to overcome the limitations of traditional low-light image enhancement methods, achieve more efficient and high-quality low-light image enhancement effects, and provide strong technical support and guarantee for many fields that rely on high-quality image data. The low-light enhancement neural network of the present invention is unsupervised (i.e., zero-reference) training, does not require paired image training, and only needs to input a dark light image and use a loss function to constrain it to train the result. Among them, the attention module can automatically learn and assign different weights to features of different channels and spatial positions, so as to integrate the channel information and spatial information of the image more comprehensively and effectively; the residual quotient is obtained by combining the low-light image and the output of the enhancement network to better adapt to the non-uniform lighting conditions commonly seen in natural low-light images, and effectively improve the model's ability to handle complex lighting changes, thereby generating more natural and visually enhanced images; at the same time, the feature assistance and accelerated training of the calibration network are combined to overcome the limitations of traditional low-light image enhancement methods, so that the network can deeply mine various feature information in low-light images, which is helpful to extract rich multi-level features. The calibration network can extract additional image feature information during the training process through its unique input convolutional layer, multiple intermediate convolutional layers and output convolutional layer structure, thus overcoming the limitations of traditional low-light image enhancement methods and achieving more efficient and high-quality low-light image enhancement effects, providing strong technical support and guarantee for many fields that rely on high-quality image data. BRIEF DESCRIPTION OF THE DRAWINGS

[0199] Figure 1 This is a specific schematic diagram of the training phase and the testing phase of the zero-reference low-light enhancement network structure for residual quotient learning proposed in the present invention.

[0200] Figure 1 (a) is a specific schematic diagram of the training phase.

[0201] Figure 1 Figure (b) is a specific schematic diagram of the test phase.

[0202] Figure 2 It is a schematic diagram of the improved residual network (i.e., ResNet module) structure of the present invention.

[0203] Figure 2 Figure (a) is a schematic diagram of the SE-RESB module structure in the ResNet module.

[0204] Figure 2 Figure (b) is a schematic diagram of the ResNet module structure.

[0205] Figure 3 It is a residual connection structure diagram of the RC-CBAM attention mechanism of the present invention.

[0206] Figure 3 Figure (a) is a schematic diagram of the RC-CBAM attention module structure.

[0207] Figure 3 Figure (b) is a schematic diagram of the channel attention module structure.

[0208] Figure 3 Figure (c) is a schematic diagram of the spatial attention module structure.

[0209] Figure 4 It is a comparison chart of the improvement of the training speed with or without using the calibration network in the present invention.

[0210] Figure 5 It is a specific structural schematic diagram of the calibration network and the enhancement network.

[0211] Figure 5 Figure (a) is a schematic diagram of the enhancement network structure.

[0212] Figure 5 Figure (b) is a schematic diagram of the calibration network structure. Specific implementation manner

[0213] The present invention will be further clarified below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications of the present invention by those skilled in the art fall within the scope defined by the appended claims of this application.

[0214] The present invention designs a zero-reference low-light image enhancement method based on residual quotient learning, as Figures 1 - 5 shown, including the following steps:

[0215] Step 1: Obtain a low-light image and preprocess it as a training set;

[0216] In the embodiment, the selected training data set is specifically a mixed data set of LSRW and LOL. A total of 1000 pictures are randomly selected from the data set, and both the input and output networks of each image adopt three channels for input. The test sets used are LSRW, LOLV1+V2, LOL SYN, MIT, MEF, Fusion, DICM, VV, and NPE for public use.

[0217] Step 2: During image training, the enhancement network, residual quotient, and calibration network are used to enhance the input image simultaneously; the framework idea of the zero-reference low-light image enhancement method based on residual quotient learning of the present invention can be expressed as:

[0218]

[0219] k represents the output of the enhancement network, that is, the part divided by the original input image, which includes the internal network and the addition and fusion operations. x represents the output image, where and respectively represent element-wise division and element-wise summation derived from the global quotient shortcut and the global residual shortcut. Net(.) is the only internal network that needs to be further designed and optimized, such as operations like the ResNet module in the enhancement network, the Laplacian sharpening module, and multiplying the sharpened feature map fea by the learnable parameter factor.

[0220] During the image training, first, the current input image input is processed by the enhancement network, as shown in Figure 1 the training part in (a). The processing process of the enhancement network includes E-Net and the plus sign. Quotient k is the result format output by the enhancement network. The processing process of the calibration network includes C-Net and the plus sign: the plus sign is the addition and fusion operation, and the division sign is the division operation.

[0221] During image training, first, the current input image is processed by the enhancement network. As shown in Figure 2 (a) the training part, the processing process of the enhancement network includes E-Net and the addition and fusion operation (represented by ). Specifically, assuming the r-th recursion of E-Net, its current input is denoted as Z r , then this enhancement process can be expressed as and the initial input is z 1 = input. And the residual learning method is adopted, that is, the information of the residual quotient is obtained by dividing the predicted low-light image by the output of the enhancement network. Subsequently, the calibration network will further process the residual quotient. The processing process of the calibration network includes C-Net and the addition and fusion operation (represented by ).

[0222] By enhancing the synergy between the enhancement network and the calibration network, the processing of the input image input is continuously optimized in each recursion, enabling the finally output image x to effectively enhance the low-light image, improve the image quality and visibility. At the same time, this design allows the network to better adapt to different degrees of low-light degradation, enhancing the robustness and generalization ability of the model.

[0223] Optimized: The residual quotient in the present invention plays a key role in image enhancement. It is an important bridge connecting the low-light input image and the desired restored image, specifically manifested in the following aspects: Redefining the enhancement task: The residual quotient x reconstructs through the relationship between the low-light image input and the output image k of the enhancement network, making the enhancement problem transformed into the learning and prediction of the residual quotient. This redefinition provides a new perspective and solution idea for image enhancement, no longer limited to the traditional decomposition-enhancement-reconstruction process. Adapting to non-uniform illumination: Different from the common practice in existing Retinex-related methods of reducing the number of channels of the quotient from three (color) to one (grayscale), this method restores the three-channel structure. This measure can better adapt to the common non-uniform illumination in natural low-light images, effectively improving the model's ability to handle complex lighting changes, thereby generating more natural and better visual-effect enhanced images. Facilitating network learning and optimization: When designing the network, a neural network is used to learn the mapping from the low-light image to the residual quotient. Since and are visually somewhat similar, assuming there is a relatively simple and easy-to-optimize connection between them, by transforming the direct mapping into learning the residual function (i.e., ), the network training process can be further simplified, improving the performance of the overall model. In the test stage, through a simple element-wise division operation the final enhanced image can be quickly generated, significantly accelerating the inference speed and meeting the efficiency requirements in practical applications.

[0224] The processing process of the enhancement network is: illu = Enhance(input); in the formula, Enhance represents the enhancement network, which includes a residual network, an attention mechanism, and a Laplacian sharpening operation to extract features, enhance, and sharpen the input image. Then, the sharpened feature map is multiplied by the learnable parameter and added to the original input image to output the enhanced image illu.

[0225] The enhanced network includes a neural network model improved based on the residual network architecture. By introducing residual blocks and combining components such as the convolutional block attention mechanism, each residual block (SE-RESB module) includes four consecutive convolutional layers, followed by a batch normalization layer and a ReLU activation function in each layer. A SE-RESB module loops three times to weight the channels of the feature map. The attention mechanism, that is, the RC-CBAM attention module includes a channel attention module and a spatial attention module, and uses multiple residual connection operations on the channel attention and spatial attention modules to combine the weights of the channel and spatial positions. The enhanced network also includes a Laplacian sharpening module.

[0226] The enhanced network performs feature extraction, enhancement, and sharpening on the input image according to its internal structure (including operations such as residual network, attention mechanism, and Laplacian sharpening), and outputs the enhanced image illu.

[0227] The specific structure of the enhanced network is as follows, as Figure 5 shown in the enhanced network part of (a) in

[0228] First, the input image data input is processed through the ResNet module to obtain the feature map fea l , and the formula can be expressed as:

[0229] fea l = ResNet(unput);

[0230] where ResNet represents the improved residual neural network. Then, the obtained feature map fea l is sharpened through the Laplacian sharpening module, and the formula is:

[0231] fea = LaplacianSharapening(fea l );

[0232] In the formula, LaplacianSharpening represents the Laplacian sharpening module. Then, the enhanced image illu is calculated according to the following formula:

[0233] illu = fea * factor + input;

[0234] That is, it is obtained by multiplying the sharpened feature map fea by the learnable parameter factor and then adding it to the original input image input. Finally, the calculated illu is restricted in the numerical range, and its value is restricted between 0.0001 and 1 to obtain the final output image illu of the enhanced network.

[0235] Next, calculate the brightness adjustment factor r:

[0236]

[0237] In the formula, the brightness adjustment factor r is obtained by dividing by the input feature map input, and torch.clamp is used to limit the brightness adjustment factor between 0 and 1 to obtain the brightness adjustment factor r (i.e., the image r after brightness adjustment).

[0238] The brightness adjustment factor r is fed into the calibration network:

[0239] att = Calibrate(r); in the formula, Calibrate represents the calibration network.

[0240] The method of the zero-reference low-light image enhancement calibration network based on residual quotient learning is as follows: The calibration network consists of an input convolutional layer, an intermediate convolutional layer, and an output convolutional layer. As shown in (b) the calibration network in Figure 5 , the entire network operation process can be expressed by the following formula (denote the input image as the brightness adjustment factor r, the output image as att (i.e., image I C-net ), and the intermediate layer feature representations are respectively F 1 、F 2 , and the estimation of the calibration network is ).

[0241] The processing process of the calibration network is as follows:

[0242] First, the input convolutional layer performs feature extraction, and the formula is:

[0243] F 1 = ReLU(BatchNorm(Conv2d(r)));

[0244] F 2 = ReLU(BatchNorm(Conv2d(F 1 )));

[0245] F 3 = ReLU(BatchNorm(Conv2d(F 2 )));

[0246]

[0247] Wherein, Conv2d represents a two-dimensional convolution operation, BatchNorm represents a batch normalization operation, ReLU represents a rectified linear unit activation function. The first formula is the input convolution layer, the second and third formulas represent intermediate convolution layers, the fourth formula represents a subtraction operation, the fifth formula is an output convolution layer with a size of 3×3, and the sixth formula represents a residual connection addition operation, that is, adding the original input image r to for optimizing the image quality. The finally output image is att.

[0248] During the training process, the enhanced image results at each stage need to be recorded simultaneously. Initialize the loss value loss to 0. Then iterate through each loop stage, and calculate the loss value loss between the input image in[i+1] and the enhanced image in[i] (i is the loop stage, ranging from 0, 1, 2, 3) at each loop stage through the loss function object Criterion u , loss u represents the loss value corresponding to the i-th loop stage, and accumulate the loss values loss u of multiple loop stages into the total loss loss. The loss value loss i can be expressed by the formula: loss i = Criterion(in[i+1], in[i]).

[0249] During the testing process, as shown in Figure 1 (b), the initialization of the convolution layer weights and the batch normalization layer weights is the same as that in the training process. If the input image is input p , first process it through the ResNet module and the Laplacian sharpening module in the enhancement network to obtain the feature representation p. The formula is: p = LaplacianSharpening(ResNet(input p ))).

[0250] Then, add the enhanced feature p to the original input image input p : p 1 = input p + p.

[0251] In the formula, p 1 is the enhanced feature map output after combination. This operation fuses the original image information on the basis of enhancing the image, aiming to adjust the enhancement effect to make the output image more in line with the actual requirements.

[0252] Then, obtain the brightness adjustment factor p 2 in the same way as the training process. The formula is:

[0253]

[0254] Finally, the final model returns the enhanced brightness adjustment factor p 2 , and this result can be used for subsequent analysis, evaluation, or further processing.

[0255] In the embodiment, as shown in (b) of Figure 2 , denote the input image data input as x, and fea l is x 3 , the processing process of the neural network module ResNet improved based on the residual network architecture in the enhancement network is as follows:

[0256] 1) First, process the input image data x through the initial convolutional layer sequence, and the formula is:

[0257] x 1 = initial_conv(x);

[0258] Obtain the feature map x after initial processing through the initial convolutional layer sequence initial_conv 1 , and the initial convolutional layer sequence includes: a two-dimensional convolutional layer with 3 input channels (corresponding to the RGB channels of the input image), 3 output channels, a batch normalization layer with 3 input channels, a ReLU activation function, a convolutional block attention module RC-CBAM with 3 input channels for channel and spatial attention processing of the feature map. A max pooling layer with a convolutional kernel size of 3×3.

[0259] 2) Then, process the feature map x 1 through the residual block sequence SE-RESB module for three cycles, and the formula is:

[0260] a 3 = residualblocks(residualblocks(residualblocks(x 1 )));

[0261] In this formula, a 3 is the feature map after further processing, and residualblocks is the SE-RESB module, as shown in (a) of Figure 2 , which is divided into the following operations:

[0262] 1a) Save the input feature map x 1 as part of the residual connection;

[0263] 2a) Perform the following series of operations on the input feature map x 1 :

[0264] a = ReLU(BatchNorm(Conv2d(x 1 )));

[0265] Conv2d represents a two-dimensional convolution operation, BatchNorm represents a batch normalization operation, ReLU represents a rectified linear unit activation function, and a is the new feature map obtained after these processes. The image undergoes this operation loop 4 times. After a series of convolution, normalization, and activation operations above, the processed feature map a is obtained 1 Then, it is processed by the channel attention mechanism SE layer, and the formula is:

[0266] a 2 = SELayer(a 1 );

[0267] In this formula, SELayer represents the SE layer. The feature map a after being processed by the channel attention mechanism SE layer 2 is added to the saved residual connection part x 1 The formula is:

[0268] a 3 = a 2 + x 1 ;

[0269] Thus, the final output feature map a of the residual block is obtained 3 。

[0270] 3) Then, the feature map a obtained through the residual block sequence 3 is processed through the final convolution layer sequence in a secondary loop, and the formula is: x 3 = final_conv(final_conv(a 3 ));

[0271] In this formula, final_conv is the final convolution layer sequence, and x 3 is the feature map after the final convolution process. The final convolution layer sequence includes: a two-dimensional convolution layer with an input channel number of 3, an output channel number of 3, a convolution kernel size of 3, a stride of 1, and a padding of 1. A batch normalization layer, a ReLU activation function. Then, the operation of the final convolution layer final_conv is looped twice

[0272] In the embodiment, the RC-CBAM attention mechanism used in the neural network module improved based on the residual network architecture is: as shown in (a) of Figure 3 , it is mainly composed of a channel attention module and a spatial attention module, aiming to perform effective attention mechanism processing on the input feature map to improve the feature extraction and understanding ability of the neural network for data such as images. The steps are as follows:

[0273] 1) Process the input feature map z through the channel attention module Cam to obtain the feature map after channel attention adjustment. The formula can be expressed as:

[0274] z 1 = Cam(z);

[0275] 2) Process the feature map z obtained in the previous step 1 through the spatial attention module Sam. The formula is:

[0276] z 2 = Sam(z 1 );

[0277] 3) Process the feature map z through the channel attention module Cam again, and add it to the original feature map z 2 : 2 Addition:

[0278] z 3 = Cam(z 2 ) + z 2 ;

[0279] 4) Process the feature map z through the spatial attention module Sam again, and add it to the original feature map z 3 : 3 The formula is:

[0280] z 4 = Sam(z 3 ) + z 3 ;

[0281] 5) Process the feature map z through the channel attention module Cam again to obtain: 4

[0282] z 5 = Cam(z 4 );

[0283] 6) Finally, process the feature map z through the spatial attention module Sam to obtain the final output feature map: 5

[0284] z 6 = Sam(z 5 );

[0285] This is the complete processing result of the RC-CBAM module on the input feature map z. Here, Cam represents the channel attention module, and Sam represents the spatial attention module.

[0286] As shown in (b) of Figure 3 , the operation of the channel attention module Cam is as follows:

[0287] 1) Create an AdaptiveMaxPool2d layer, which pools the input feature map into a feature map y of size 1×1, expressed as: maxpool , expressed as:

[0288] y maxpool = AdaptiveMaxPool2d(1)(Q);

[0289] In the formula, Q is the input feature map, and AdaptiveMaxPool2d(1)(Q) represents the adaptive max pooling operation of the input feature map to size 1. Then, through a sequence consisting of multiple fully connected layers, including:

[0290] i. The first fully connected layer Linear expands the number of input channels. The formula is:

[0291] Q 1 = Linear(inChannel, int(inChannel * r))(y maxpool );

[0292] In the formula, inChannel represents the number of input channels, y maxpool is the input after being flattened by the max pooling layer, Linear represents the fully connected layer operation, and int(inChannel * r) represents the operation of expanding the number of input channels by r times.

[0293] ii. Then, it is processed by the hyperbolic tangent activation function Tanh. The formula is:

[0294] Q 2 = Tanh(Q 1 );

[0295] iii. The second fully connected layer Linear restores the dimension of the feature map Q 2 to the original input size. The formula is:

[0296] Q 3 = Linear(int(inChannel * r), inChannel)(Q 2 );

[0297] iv. Finally, it is processed by the Sigmoid activation function to obtain the weight weight of the max pooling branch max , and the formula is:

[0298] weight max = Sigmoid(Q 3 );

[0299] 2) Similarly, repeat the above operations to create an adaptive average pooling layer. The formula is:

[0300] y avgpool = AdaptiveAvgPool2d(1)(Q);

[0301] AdaptiveAvgPool2d(1)(Q) represents an adaptive average pooling operation that reduces the input feature map Q to a size of 1.

[0302] Create a sequence of fully connected layers similar to the max pooling layer sequence to process the average pooling branch y avgpool , obtaining the weight weight of the average pooling branch avg , and the processing formula for each layer is similar to that of the above max pooling branch. The specific processing process is as follows:

[0303] i. The first fully connected layer Linear expands the number of input channels, and the formula is:

[0304] q 1 = Linear(inChannel, int(inChannel * r))(y avgpool );

[0305] In the formula, inChannel represents the number of input channels, y avgpool is the input after being flattened by the average pooling layer, Linear represents the fully connected layer operation, and int(inChannel * r) represents the operation of expanding the number of input channels by r times;

[0306] ii. Then, it is processed by the hyperbolic tangent activation function Tanh, and the formula is:

[0307] q 2 = Tanh(q 1 );

[0308] iii. The second fully connected layer Linear restores the dimension of the feature map q 2 to the original input size, and the formula is:

[0309] q 3 = Linear(int(inChannel * r), inChannel)(q 2 );

[0310] iv. Finally, it is processed by the Sigmoid activation function to obtain the weight weight of the average pooling branch avg , and the formula is:

[0311] weight avg = Sigmoid(q 3 );

[0312] 3) Add the maximum pooling branch weight weight max and the average pooling branch weight weight avg to obtain the fused weight weight:

[0313] weight = weight max + weight avg .

[0314] Reshape the fused weight weight into a shape matching the input feature map Q. Assuming the shape of the weight tensor is (h, w), the reshape formula is:

[0315] Mc = reshape(weight, (h, w, 1, 1)).

[0316] In the formula, reshape is the reshape operation, Mc is the reshaped weight, and the tensor of weight is reshaped from a two-dimensional tensor to a new four-dimensional tensor (h, w, 1, 1), where: h and w represent the height and width of the new tensor respectively. 1 and 1 indicate that the last two dimensions of the new tensor are single-element dimensions. weight is a [3, 3] matrix, and reshape(weight, (3, 3, 1, 1)) will change it into a four-dimensional tensor of [3, 3, 1, 1], which can be used for convolution operations with single-channel inputs.

[0317] Finally, multiply the reshaped weight Mc element-wise with the input feature map Q to obtain the feature map z after channel attention adjustment i , and the formula is: z i = Mc * Q.

[0318] As shown in (c) of Figure 3 , the spatial attention module operates as follows:

[0319] First, take the maximum value (through the torch.max operation) and the average value (through the torch.mean operation) of the input feature map l in the channel dimension respectively, and then concatenate these two results in the channel dimension (torch.cat) to obtain a new feature map m. The formula is: m = torch.cat((max(l), mean(t)).

[0320] In this formula, torch.cat is the concatenation operation, max is the operation of taking the maximum value of feature map l, and mean is the operation of taking the average value of feature map l. The new feature map m then passes through a convolutional layer with a kernel size of 7 and a padding of 3 and the Sigmoid function to obtain the weight Ms at each spatial position, which is used to reflect the importance of features at different spatial positions. Finally, the obtained spatial weight Ms is multiplied element-wise with the input feature map l to obtain the feature map t adjusted by spatial attention, and the formula is as follows: t = Ms * l.

[0321] This process realizes the reweighting of the input feature map according to the importance of spatial positions, enabling the network to focus on the key spatial regions in the image.

[0322] Step 3: Train the constructed low-light enhancement neural network with the training set to obtain a trained low-light enhancement neural network;

[0323] Step 4: Load the enhancement network with the pre-trained model obtained in the image test to process the input image to obtain the final picture.

[0324] In another embodiment, a zero-reference low-light image enhancement method based on residual quotient learning is mentioned, including an enhancement network and a calibration network. Among them:

[0325] The enhancement network is used to perform feature extraction, feature weighting, detail enhancement, etc. on the image, and finally outputs a high-quality enhanced image;

[0326] The enhancement network contains a neural network model improved based on the residual network architecture, which enables the model to learn residual features, effectively alleviates the problem of gradient disappearance during the training of deep networks. Different levels of residual blocks can extract multi-level features from low-level to high-level, from detail features such as edges and textures to semantic features, providing a rich feature basis for subsequent image enhancement. While maintaining details, the enhanced image can better reflect the semantic content of the image. It also contains a Laplacian sharpening module, which performs filtering and image enhancement sharpening operations on the image, traversing the batch, height, and width dimensions through nested loops.

[0327] The neural network model improved based on the residual network architecture contains an RC-CBAM attention mechanism module, which is used to organically combine the weights of channels and spatial positions, realizing comprehensive and accurate weighting of image features. This synergistic effect enables the network to focus on the key information in the image simultaneously from both the channel and spatial dimensions. During the low-light image enhancement process, it can more effectively extract and utilize useful features such as weak object edges and dim texture details.

[0328] The residual quotient has advantages in the enhancement processing of low - illumination images with uneven illumination. It can better adapt to the common non - uniform illumination situations in natural low - illumination images, effectively improve the processing ability for complex illumination changes, and thus generate more natural and better - visual - effect enhanced images.

[0329] The calibration network is used to provide additional feature information support for the enhancement network, which helps to accelerate the convergence speed of the enhancement network. By providing additional gradient information, the brightness and contrast can be adjusted more reasonably when enhancing the image, making the enhancement effect more natural and accurate.

[0330] In another embodiment, a computer system is provided, including a memory and a processor. The memory is used to store computer programs / instructions; the processor is used to execute the computer programs / instructions to implement the zero - reference low - illumination image enhancement method based on residual quotient learning.

[0331] In the embodiment, we propose a new improved residual network, which includes our newly proposed attention combination mechanism RC - CBAM, used to organically combine the weights of channels and spatial positions, and achieve comprehensive and accurate weighting of image features. This synergistic effect enables the network to focus on the key information in the image from both the channel and spatial dimensions simultaneously.

[0332] To prove the effectiveness of the proposed module, we designed several ablation experiments.

[0333] As shown in Table 1, the combination of the SE - RESB module and the RC - CBAM module has the best image enhancement effect, with SSIM being 0.9152, PSNR being 20.87, and LPIPS being 0.1069, all of which are the highest. The combination of the SE - RESB module and CBAM is not as good as the combination with the RC - CBAM module, proving that our module is more helpful for the fusion of multi - attention modules compared to the CBAM module.

[0334] Table 1. Data comparison table of combination experiments of different attention modules:

[0335] SE-RESB module CBAM module RC-CBAM module PSNR↑ SSIM↑ LPIPS↓ × × × 19.45 0.8927 0.1306 √ × × 20.81 0.9114 0.1139 √ √ × 16.99 0.8596 0.1657 × × √ 20.52 0.9090 0.1138 × √ × 18.15 0.8432 0.2183 √ × √ 20.87 0.9152 0.1069

[0336] We also designed the influence of the structure of the SE - RESB module on the image enhancement effect. The results are shown in Table 2. It can be seen that when the number of modules is 3, the number of convolutional layers is 4, and the number of channels is 3, the SE - RESB module has the best effect.

[0337] Table 2. Influence table of the structure of the SE - RESB module on the image enhancement effect:

[0338]

[0339] In addition, we designed ablation experiments on the loss function. We used the designed enhancement network and calibration network, and trained with different combinations of loss functions for the same number of times to compare the rationality and effectiveness of the loss function settings.

[0340] As shown in Table 3, it can be seen that the combination of fidelity loss, smoothness loss, color loss, and perceptual loss is effective. Compared with other combinations among them, the results of their combined training can reach SSIM of 0.9152, PSNR of 20.87, and LPIPS of 0.1069. All results are the best, indicating that the loss function we designed has good training performance.

[0341] Table 3: Ablation experiment table of different combinations of loss functions:

[0342]

[0343]

[0344] In addition, we designed ablation experiments on the calibration network. We used the designed enhancement network and calibration network and divided them into two groups for training. One group was the trend of the loss function value changing with the increase in the number of training times without the calibration network, and the other group was the trend of the loss function value changing with the increase in the number of training times with the calibration network. The results are as Figure 4 shown.

[0345] It can be seen that the loss function value with the calibration network decreases faster with the increase in the number of training times, and the training speed is improved. Compared with the loss function value without the calibration network, it can achieve better training results.

[0346] A zero-reference low-light image enhancement method based on residual quotient learning, comprising the following steps:

[0347] Step 1: Obtain a low-light image and preprocess it as a training set;

[0348] Step 2: During image training, simultaneously use the enhancement network and the calibration network to enhance the input image;

[0349] The enhanced network includes a neural network model improved based on the residual network architecture. By introducing residual blocks and combining components such as convolutional block attention mechanism, each residual block includes four consecutive convolutional layers, followed by a batch normalization layer and a ReLU activation function for each convolutional layer, and an SE-RESB module for weighting the channels of the feature map. The attention mechanism consists of an SE-RESB module and an RC-CBAM attention module. The RC-CBAM attention module includes a channel attention sub-module and a spatial attention sub-module, and multiple residual connection operations for the channel attention sub-module and the spatial attention sub-module are used to combine the weights of the channels and spatial positions. The enhanced network also includes a Laplacian sharpening module.

[0350] The calibration network includes an input convolutional layer and 3 convolutional blocks, and each convolutional block includes consecutive convolutional layers, a batch normalization layer, and an output convolutional layer.

[0351] During the training process, the enhanced network and the calibration network share the input image data. The enhanced network first processes the input image to extract additional feature information, and then passes this information to the calibration network. Based on its own structure and algorithm, the calibration network combines the features provided by the enhanced network to further enhance the image and calculate the loss value.

[0352] Step 3: Train the constructed low-light enhancement neural network with the training set to obtain a trained low-light enhancement neural network.

[0353] Step 4: Load the enhanced network with the pre-trained model obtained in the image test to process the input image to obtain the final picture.

[0354] For the above zero-reference low-light image enhancement method, the framework of the zero-reference low-light image enhancement method based on residual quotient learning is expressed as:

[0355]

[0356] q represents the illuminance, y and x are the degraded observation value and the expected recovery value respectively, where and represent element-wise division and element-wise summation derived from the global quotient shortcut and the global residual shortcut respectively. Net(.) is the internal network that only needs to be further designed and optimized, and SystemOutput represents the output picture of the system network.

[0357] For the above zero-reference low-light image enhancement method, in the enhanced network, the network method Net of the enhancement framework is:

[0358] First, the input image data input is processed through the ResNet module to obtain the feature map fea, and the formula can be expressed as:

[0359] fea = ResNet(input);

[0360] Where ResNet represents an improved residual neural network; then, the obtained feature map fea is sharpened through the Laplacian sharpening module, and the formula is:

[0361] fea = LaplacianSharpening(fea);

[0362] In the formula, LaplacianSharpening represents the Laplacian sharpening module; then, the enhanced image illu is calculated according to the following formula:

[0363] illu = fea * factor + input;

[0364] That is, it is obtained by multiplying the sharpened feature map fea by the learnable parameter factor and then adding it to the original input image input; finally, the calculated illu is restricted in the numerical range, and its value is restricted between 0.0001 and 1 to obtain the final output image of the enhancement network and return it;

[0365] The method for calibrating the network is as follows: The composition of the calibration network is divided into input, intermediate, and output convolutional layers; the operation process of the entire network can be expressed by the following formula. Assuming the input image is I input , the output image is I C-net , and the feature representations of each intermediate layer are F i , where i represents different processing stages, and the noise estimation is expressed as First, the input convolutional layer performs feature extraction, and the formula is:

[0366] F 1 = ReLU(BatchNorm(Conv2d(I input )));

[0367] In the formula, Conv2d represents a two-dimensional convolution operation, BatchNorm represents a batch normalization operation, and ReLU represents a rectified linear unit activation function. This formula represents the feature output after the input image passes through the input convolutional layer. Then, the image is processed multiple times by the intermediate convolutional layer group. For the i-th loop, the formula is:

[0368] F i = F i-1 + I input

[0369] Fi = ReLU(BatchNorm(Conv2d(F i-1 )));

[0370] Among them, the first formula reflects the residual connection operation, and the second formula represents the processing process of the feature through the sub-structure in the intermediate convolutional layer group, continuously updating the feature representation; finally, after the conversion of the output convolutional layer and the calibration network calculation, the formula is:

[0371]

[0372] Among them, the first formula means that the feature after multiple processes passes through the output convolutional layer to obtain the image enhancement estimation, and the second formula calculates the final image according to the calibration network, that is, the original input image minus the enhanced estimation image.

[0373] In the above zero-reference low-light image enhancement method, in the enhancement network, the neural network module improved based on the residual network architecture is:

[0374] First, the input image data x is processed through the initial convolutional layer sequence, and the formula is:

[0375] x 1 = initial_Conv(x);

[0376] The initial convolutional layer sequence initial_Conv is used to obtain the feature map x after initial processing 1 , and the initial convolutional layer sequence includes: a two-dimensional convolutional layer with an input channel number of 3, the input channel number corresponding to the RGB channels of the input image, an output channel number of 3, a batch normalization layer with an input channel number of 3, a ReLU activation function, a convolutional block attention module RC-CBAM with an input channel number of 3 for channel and spatial attention processing of the feature map; a max pooling layer with a convolutional kernel size of 3;

[0377] Then, the feature map x 1 is processed through the residual block sequence, and the formula is:

[0378] x 2 = residualblocks(x 1 );

[0379] In this formula, x 2 is the feature map after further processing;

[0380] residualblocks is the residual block sequence, which is divided into the following operations:

[0381] (1), Save the input feature map a as part of the residual connection residual;

[0382] (2) Perform the following series of operations on the input feature map a:

[0383] After being processed by the first convolutional layer, the formula is:

[0384] a 1 = Conv2d 1 (a);

[0385] In this formula, Conv2d 1 (a) represents a convolutional layer with a kernel size of 3, a stride of 1, and a padding of 1;

[0386] (3) Then, it is processed by the corresponding batch normalization layer, and the formula is:

[0387] a 2 = BatchNorm2d 1 (a 1 );

[0388] (4) Then, it is processed by the ReLU activation function, and the formula is:

[0389] a 3 = ReLU(a 2 );

[0390] (5) Then, the image passes through the first convolutional layer operation three times. After a series of convolutional, normalization, and activation operations, the obtained feature map a 4 is then processed by the channel attention SE-RESB module, and the formula is:

[0391] a 5 = SELayer(a 4 );

[0392] In this formula, SELayer represents the SE-RESB module; the feature map a after being processed by the channel attention SE-RESB module 5 is added to the saved residual connection part residual, and the formula is:

[0393] a 6 = a 5 + residual;

[0394] Thus, the final output feature map a of the residual block is obtained 6 and returned;

[0395] Then, the feature map x obtained through the residual block sequence 2 is processed through the final convolutional layer sequence, and the formula is:

[0396] x 3= final_conv(x 2 );

[0397] In this formula, final_conv is the final convolutional layer sequence, and x 3 is the feature map after final convolutional processing; the final convolutional layer sequence includes: a two-dimensional convolutional layer with an input channel number of 3, an output channel number of 3, a convolutional kernel size of 3, a stride of 1, and a padding of 1; a batch normalization layer, a ReLU activation function, and an adaptive average pooling layer that pools the feature map into a feature map of size 1×1, and finally passes through a two-dimensional convolutional layer with an output channel number of 3 to convert the feature map back to the RGB channel number form.

[0398] For the above zero-reference low-light image enhancement method, the RC-CBAM attention mechanism used in the neural network module improved based on the residual network architecture is mainly composed of a channel attention module and a spatial attention module, and the steps are as follows:

[0399] (1) Process the input feature map z through the channel attention module Cam to obtain the feature map after channel attention adjustment. The formula can be expressed as:

[0400] x 1 = Cam(x);

[0401] (2) Process the feature map x obtained in the previous step 1 through the spatial attention module Sam. The formula is:

[0402] z 2 = Sam(z 1 );

[0403] (3) Process the feature map z again 2 through the channel attention module Cam and add it to the original feature map z 2 :

[0404] z 3 = Cam(z 2 ) + z 2 ;

[0405] (4) Process the feature map z again 3 through the spatial attention module Sam and add it to the original feature map x 3 : The formula is:

[0406] z 4 = Sam(z 3 ) + z 3 ;

[0407] (5) Process the feature map z again 4Processed by the channel attention module Cam to obtain:

[0408] z 5 = Cam(z 4 );

[0409] (6) Finally, the feature map z 5 is processed by the spatial attention module Sam to obtain the final output feature map:

[0410] z 6 = Sam(z 5 );

[0411] This is the complete processing result of the RC-CBAM module on the input feature map z and is returned; where Cam represents the channel attention module and Sam represents the spatial attention module;

[0412] The operation of the channel attention module is as follows:

[0413] (1) Create an adaptive max pooling layer AdaptiveMaxPool2d, which pools the input feature map into a feature map y of size 1×1 maxpool , expressed as:

[0414] y maxpool = AdaptiveMaxPool2d(1)(y);

[0415] In the formula, y is the input feature map, and AdaptiveMaxPool2d(1)(y) represents the adaptive max pooling operation of the input feature map to size 1; then through a sequence composed of multiple fully connected layers, including:

[0416] i. The first fully connected layer Linear expands the number of input channels, and the formula is:

[0417] y 1 = Linear(inChannel, int(inChannel * r))(y maxpool );

[0418] In the formula, inChannel represents the number of input channels, y maxpool is the input after being flattened by the max pooling layer, representing the fully connected layer operation, and int(inChannel * r) represents the operation of expanding the number of input channels by r times;

[0419] ii. Then, it is processed by the hyperbolic tangent activation function Tanh, and the formula is:

[0420] y 2 = Tanh(y 1 );

[0421] iii. The second fully connected layer Linear transforms the feature map y 2 back to the original input size, and the formula is:

[0422] y 3 = Linear(int(inChannel * r), inChannel)(y 2 );

[0423] iv. Finally, after being processed by the Sigmoid activation function, the weight weight of the max pooling branch is obtained max , and the formula is:

[0424] weight max = Sigmoid(y 3 );

[0425] v. Similarly, repeat the above operations to create an adaptive average pooling layer, and the formula is:

[0426] y avgpool = AdaptiveAvgPool2d(1)(y);

[0427] (2). Create a sequence of fully connected layers similar to the max pooling layer sequence to process the average pooling branch y avgpool , and obtain the weight weight of the average pooling branch avg , and the processing formulas of each layer are similar to those of the above max pooling branch;

[0428] i. The first fully connected layer Linear expands the number of input channels, and the formula is:

[0429] q 1 = Linear(inChannel, int(inChannel * r))(y avgpool );

[0430] In the formula, inChannel represents the number of input channels, and y avgpool is the input after being flattened by the average pooling layer, representing the fully connected layer operation, and int(inChannel * r) represents the operation of expanding the number of input channels by r times;

[0431] ii. Then, it is processed by the hyperbolic tangent activation function Tanh, and the formula is:

[0432] q 2 = Tanh(q 1 );

[0433] iii. The second fully connected layer Linear transforms the feature map q 2 back to the original input size, and the formula is:

[0434] q 3 = Linear(int(inChannel * r), inChannel)(q 2 );

[0435] iv. Finally, after being processed by the Sigmoid activation function, the weight weight of the max pooling branch is obtained avg , and the formula is: weight avg = Sigmoid(q 3 );

[0436] v. Add the weight weight of the max pooling branch and the weight weight of the average pooling branch to obtain the fused weight weight: max and the weight weight of the average pooling branch avg weight = weight

[0437] + weight max avg ;

[0438] Reshape the fused weight weight into a shape matching the input feature map y. Assuming the shape of the weight tensor is (h, w), the reshape formula is:

[0439] Mc = reshape(weight, (h, w, 1, 1));

[0440] In the formula, reshape is the reshape operation, and Mc is the reshaped weight. Finally, multiply the reshaped weight Mc element-wise with the input feature map y to obtain the feature map z after channel attention adjustment. The formula is:

[0441] z = Mc * y;

[0442] (3). The operations of the spatial attention module are as follows: First, take the maximum value and the average value of the input feature map y in the channel dimension respectively. The maximum value is obtained through the torch.max operation, and the average value is obtained through the torch.mean operation. Then, concatenate these two results in the channel dimension, that is, torch.cat, to obtain a new feature map m. The formula is:

[0443] m = torch.cat((max(y), mean(y));

[0444] ​In this formula, torch.cat is the concatenation operation, max is the operation of taking the maximum value of the feature map y, and mean is the operation of taking the average value of the feature map y; then, a convolutional layer with a kernel size of 7 and a padding of 3 and a Sigmoid function are used to obtain the weights at each spatial position, which are used to reflect the importance of features at different spatial positions; finally, the obtained spatial weight Ms is multiplied element-wise with the input feature map m to obtain the feature map t adjusted by spatial attention, and the formula is as follows:

[0445] t = Ms * m.

[0446] In this example, the neural network, residual network, attention mechanism, SE layer, and Laplacian sharpening module are all professional terms in the field, which are prior arts and not the main improvement points of the present invention, so they will not be elaborated here.

[0447] In this example, a zero-reference low-light image enhancement method based on residual quotient learning is designed, which includes a neural network architecture of an enhancement network and a calibration network. The enhancement network first processes the input image to extract additional feature information, then calculates the residual quotient, and the residual quotient is sent to the calibration network. The calibration network, based on its own structure and algorithm, combines the features provided by the residual quotient to further enhance the image and calculate the loss value. The enhancement network and the calibration network continuously optimize their own performance through a collaborative parameter update mechanism to improve the overall network's ability to enhance low-light images, so that the generated enhanced images can achieve better effects in terms of visual quality, detail retention, semantic information, etc. This technology can also effectively improve the brightness, contrast, detail clarity, and semantic information retention of low-light images, achieve a good balance between the number of parameters and performance, be applicable to multiple fields, and have broad application prospects and significant practical value.

Claims

1. A zero-reference low-light image enhancement method based on residual quotient learning, characterized in that: The following steps are involved: Step 1: Obtain low-light images and preprocess them as training sets; Step 2: construct a low-light enhancement neural network based on a residual quotient learning framework, which includes an enhancement network and a calibration network; Step 3: Train the constructed low-light enhancement neural network based on the residual quotient learning framework through the training set to obtain a trained low-light enhancement neural network model; Step 4: Use the enhanced network in the trained low-light enhancement neural network model to process the input low-light image to obtain the final image.

2. The zero-reference low-light image enhancement method based on residual quotient learning according to claim 1, characterized in that: In step 2: The processing of the enhanced network is: illu=Enhance(input); Among them, Enhance represents the enhanced network, input is the input image of the enhanced network, and illu is the output image of the enhanced network; The processing process of the residual quotient is: Calculate the brightness adjustment factor r by enhancing the output image illu of the network: torch.clamp(r,0,1) means limiting the brightness adjustment factor r between 0 and 1 to obtain the brightness adjustment factor r; the brightness adjustment factor r is the residual quotient; The process of calibrating the network is: Send the brightness adjustment factor r to the calibration network; att = Calibrate(r); In the formula, Calibrate represents the calibration network, and att is the output image of the calibration network.

3. The zero-reference low-light image enhancement method based on residual quotient learning according to claim 2, characterized in that: In the step 2, the specific processing process of enhancing the network is: Step 2a: The input image data is processed by the ResNet module to obtain the feature map fea l , the formula is: fea l =ResNet(input); Among them, ResNet represents the improved residual neural network; Step 2b: Get the feature map fea l The sharpening process is performed through the Laplace sharpening module, and the formula is: fea=LaplacianSharpening(fea l ); In the formula, LaplacianSharpening represents the Laplacian sharpening module; Step 2c, calculate the enhanced image illu: illu=fea*factor+input; That is, the enhanced image illu is obtained by multiplying the sharpened feature map fea with the learnable parameter factor and then adding it to the original input image input; Step 2d: Limit the numerical range of the calculated illu to between 0.0001 and 1 to obtain the final output image of the enhanced network.

4. The zero-reference low-light image enhancement method based on residual quotient learning according to claim 3, characterized in that: In the step 2a, the processing process of the improved residual neural network ResNet module is: Step 2a1: If the input image data is denoted as x, the input image data x is processed by the initial convolution layer sequence initial_conv, and the formula is: x1 = initial_Conv(x); The initial convolutional layer sequence initial_conv includes a 2D convolutional layer, a batch normalization layer, a ReLU activation function, a convolutional block attention module RC-CBAM, and a maximum pooling layer with a convolution kernel size of 3×3; Step 2a2: Process the feature map x1 through the residual block sequence SE-RESB module for three cycles to obtain the processed feature map a3. The formula is: a3=residualblocks(residualblocks(residualblocks(x1))); In this formula, residualblocks is the SE-RESB module; Step 2a3: The feature map a3 obtained by the residual block sequence SE-RESB module is processed twice through the final convolution layer sequence final_conv. The formula is: x3=final_conv(final_conv(a3)); In the formula, final_conv is the final convolutional layer sequence, and x3 is the feature map obtained after the final convolutional layer sequence final_conv is processed twice; The final convolutional layer sequence final_conv consists of a 2D convolutional layer, a batch normalization layer, and a ReLU activation function.

5. The zero-reference low-light image enhancement method based on residual quotient learning according to claim 4, characterized in that: In step 2a2, the processing process of the SE-RESB module is: Step 2a21, save the input feature map x1 as part of the residual connection; Step 2a22: Perform the following operations on the input feature map x1: a=ReLU(BatchNorm(Conv2d(x1))); Conv2d represents a two-dimensional convolution operation, BatchNorm represents a batch normalization operation, ReLU represents a rectified linear unit activation function, and a represents a new feature map obtained after processing; Step 2a23, the characteristic graph a is recycled three times through step 2a22 to obtain the processed characteristic graph a1; The feature map a1 is then processed by the channel attention mechanism SE layer, and the formula is: a2 = SELayer(a1); In the formula, SELayer represents the SE layer; Add the feature map a2 processed by the channel attention mechanism SE layer to the residual connection part x1 saved in step 2a21. The formula is: a3=a2+x1; Thus, the final output feature map a3 of the SE-RESB module is obtained.

6. The zero-reference low-light image enhancement method based on residual quotient learning according to claim 4, characterized in that: In step 2a1, the processing process of the convolutional block attention module RC-CBAM is: Step 2a11, let z be the input feature map of the convolutional block attention module RC-CBAM; The input feature map z is processed by the channel attention module Cam to obtain the feature map z1 after channel attention adjustment. The formula is: z1=Cam(z); Step 2a12: Process the feature map z1 obtained in the previous step through the spatial attention module Sam. The formula is: z2=Sam(z1); Step 2a13, process the feature map z2 again through the channel attention module Cam and add it to the original feature map z2. The formula is: z3=Cam(z2)+z2; Step 2a14, the feature map z3 is processed by the spatial attention module Sam and added to the original feature map z3. The formula is: z4=Sam(z3)+z3; Step 2a15, again process the feature map z4 through the channel attention module Cam to obtain: z5=Cam(z4); Step 2a16. Finally, the feature map z5 is processed by the spatial attention module Sam to obtain the final output feature map z6: z6=Sam(z5).

7. The zero-reference low-light image enhancement method based on residual quotient learning according to claim 6, characterized in that: In step 2a1, the processing process of the channel attention module Cam is: (1) Create an adaptive maximum pooling layer AdaptiveMaxPool2d, which pools the input feature map Q into a feature map y of size 1×1 maxpool , the formula is: y maxpool =AdaptiveMaxPool2d(1)(Q); In the formula, Q is the input feature map, and AdaptiveMaxPool2d(1)(Q) represents the adaptive maximum pooling operation of Q-ing the input feature map to a size of 1; It then passes through a sequence of multiple fully connected layers, including: (1a) The first fully connected layer Linear expands the number of input channels to obtain Q1, the formula is: Q1=Linear(inChannel,int(inChannel*r))(y maxpool ); In the formula, inChannel represents the number of input channels, y maxpool It is the input after flattening by the maximum pooling layer, Linear represents the fully connected layer operation, and int(inChannel*r) represents the operation of expanding the number of input channels by r times; (1b), then, after being processed by the hyperbolic tangent activation function Tanh, the formula is: Q2 = Tanh(Q1); (1c) The second fully connected layer Linear converts the feature map Q2 dimension back to the original input size. The formula is: Q3=Linear(int(inChannel*r),inChannel)(Q2); (1d), and finally processed by the Sigmoid activation function to obtain the weight of the maximum pooling branch max , the formula is: weight max =Sigmoid(Q3); (2) Create an adaptive average pooling layer AdaptiveAvgPool2d, the formula is: y avgpool =AdaptiveAvgPool2d(1)(Q); AdaptiveAvgPool2d(1)(Q) represents the adaptive average pooling operation of Q-ing the input feature map to size 1; It then passes through a sequence of multiple fully connected layers, including: (2a) The first fully connected layer Linear expands the number of input channels. The formula is: q1=Linear(inChannel,int(inChannel*r))(y avgpool ); In the formula, inchannel represents the number of input channels, y avgpool It is the input after flattening by the average pooling layer, Linear represents the fully connected layer operation, and int(inChnnel*r) represents the operation of expanding the number of input channels by r times; (2b) Then, after being processed by the hyperbolic tangent activation function Tanh, the formula is: q2 = Tanh(q1); (2c) The second fully connected layer Linear converts the dimension of the feature map q2 back to the original input size. The formula is: q3=Linear(int(inChannel*r),inChannel)(q2); (2d) Finally, after Sigmoid activation function processing, the weight of the average pooling branch is obtained. avg , the formula is: weight avg =Sigmoid(q3); (3) Weight the maximum pooling branch max and average pooling branch weight avg Add them together to get the fused weight, the formula is: weight=weight max +weight avg ; (4) Reshape the fused weights to match the input feature map Q. Assuming the shape of the weight tensor is (h, w), the reshaping formula is: Mc=reshape(weight,(h,w,1,1)); In the formula, reshape is the reshaping operation, Mc is the reshaped weight, and the weight tensor is reshaped from a two-dimensional tensor to a new four-dimensional tensor (h,w,1,1), where: h and w represent the height and width of the new tensor respectively; 1 and 1 indicate that the last two dimensions of the new tensor are single-element dimensions; Finally, the reshaped weight Mc is multiplied element-by-element by the input feature map Q to obtain the feature map z after channel attention adjustment. i , the formula is: z i =Mc*Q。 8. The zero-reference low-light image enhancement method based on residual quotient learning according to claim 6, characterized in that: In step 2a1, the processing process of the spatial attention module Sam is: Take the maximum value and average value of the input feature map l in the channel dimension, and then concatenate the two results in the channel dimension to obtain a new feature map m. The formula is: m=torch.cat((max(l),mean(l)); In this formula, torch.cat is a concatenation operation, max(l) is an operation to obtain the maximum value of feature map l, and mean(l) is an operation to obtain the average value of feature map l. Next, the new feature map m is passed through a convolutional layer and a Sigmoid function to obtain the weight of each spatial position; Finally, the obtained spatial weight Ms is multiplied element-by-element by the input feature map l to obtain the feature map t after spatial attention adjustment. The formula is: t=Ms*l.

9. The zero-reference low-light image enhancement method based on residual quotient learning according to claim 2, characterized in that: The processing process of the calibration network is as follows: The input image is recorded as the brightness adjustment factor r, and the output image is recorded as att; First, the brightness adjustment factor r is input into the input convolution layer for feature extraction to obtain the feature map F1. The formula is: F1=ReLU(BatchNorm(Conv2d(r))); Conv2d represents a two-dimensional convolution operation, BatchNorm represents a batch normalization operation, and ReLU represents a rectified linear unit activation function; The feature map F1 is input into the middle convolution layer and then feature extraction is performed to obtain the feature map F2. The feature map F2 is input into the second middle convolution layer and then feature extraction is performed to obtain the feature map F3. The formula is: F2=ReLU(BatchNorm(Conv2d(F1))); F3=ReLU(BatchNorm(Conv2d(F2))); Then subtract the brightness adjustment factor r from the feature map F3 to obtain the feature map The formula is: The feature map After the output convolution layer operation, the feature map is obtained The formula is: The feature map The residual connection and addition operation are performed with the brightness adjustment factor r, and the formula is:

Citation Information

Cited By

  • Low-illumination image dynamic enhancement method and device, medium and program product

    CN121073804A