Wavelet mamba low-light image enhancement method based on illumination prior

By employing a wavelet Mamba low-light image enhancement method based on illumination priors, and utilizing an illumination estimation module and a wavelet Mamba encoding/decoding network, the problems of uneven illumination distribution and noise sensitivity in existing methods are solved, achieving efficient image enhancement and lightweight processing.

CN120912473BActive Publication Date: 2026-05-08SHANGHAI MUNA INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI MUNA INFORMATION TECHNOLOGY CO LTD
Filing Date
2025-08-25
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing deep learning-based low-light image enhancement methods lack long-range dependency, resulting in uneven illumination distribution, overexposure and underexposure in the enhanced images, difficulty in capturing global structure and semantic information, and sensitivity to noise, which affects image quality.

Method used

We employ a wavelet Mamba low-light image enhancement method based on illumination priors. By constructing an illumination estimation module and a wavelet Mamba encoding/decoding network, we utilize the quaternary illumination prior information to process low-frequency and high-frequency information in the frequency domain. Combined with the illumination prior fusion module, we improve the global enhancement effect of the image and suppress noise.

Benefits of technology

It effectively improves illumination distribution, enhances detail recovery capabilities, reduces computational resource consumption, improves image quality, maintains naturalness and detail, and adapts to image recovery under different lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912473B_ABST
    Figure CN120912473B_ABST
Patent Text Reader

Abstract

The application discloses a wavelet Mamba low-light image enhancement method based on illumination prior, and specific steps are as follows: step 1, an illumination estimation module is constructed, a low-light image is taken as input, and illumination quaternion prior is obtained; step 2, a wavelet Mamba encoding and decoding network is constructed, the low-light image and the illumination quaternion prior are input into the wavelet Mamba encoding and decoding network, and an enhanced image is obtained; step 3, a whole network composed of the illumination estimation module and the wavelet Mamba encoding and decoding network is trained, and a trained whole network is obtained; and step 4, a low-light image to be enhanced is input into the trained whole network, and a final enhanced image is obtained. The method can enhance the image while maintaining naturalness and details, and the enhanced image has high quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image enhancement methods, specifically relating to a wavelet Mamba low-light image enhancement method based on illumination prior. Background Technology

[0002] Low-light enhancement technologies have wide applications across various fields, including visual surveillance, autonomous driving, and computational photography. In particular, smartphone photography has become ubiquitous and prominent. However, taking photos with a smartphone camera in dimly lit environments is especially challenging due to limitations in camera aperture, real-time processing requirements, and memory.

[0003] In recent years, with the rapid development of deep learning technology, low-light image enhancement methods based on deep learning have developed rapidly. Deep learning methods can adapt to different scenes and lighting conditions, effectively process complex and variable real-world image data, and provide more accurate and comprehensive prior information on lighting for subsequent image enhancement processing. This significantly improves the image enhancement effect and adaptability, and expands the application scope and practical value of low-light image enhancement technology.

[0004] Learning-based models, including CNN-based and Transformer-based methods, primarily construct convolutional neural networks (CNNs) or variant networks, using a large number of low-light images and their corresponding normal-light images as training data. The network automatically learns the mapping relationship between low-light and normal-light images, thereby generating clearer, more natural, and more detailed enhanced images. However, obtaining paired training data is often difficult in practical applications, which limits the model's generalization ability to some extent. Although deep learning-based methods have achieved significant results in low-light image enhancement, some problems and challenges remain. Deep learning-based methods lack long-range dependencies, leading to uneven illumination distribution, overexposure and underexposure in enhanced images, affecting image quality. Furthermore, most deep learning-based methods process in the raw space, mainly focusing on local pixel values ​​and texture information, making it difficult to capture global structure and semantic information. Their ability to distinguish and process different regions in complex scenes is limited, and they are sensitive to noise, easily amplifying noise during enhancement, affecting image quality. Therefore, there is still a need for a low-light image enhancement method that can improve global enhancement effects, reduce computational resource consumption, and effectively suppress noise. Summary of the Invention

[0005] The purpose of this invention is to provide a wavelet Mamba low-light image enhancement method based on illumination prior, which solves the problem of poor image quality in existing methods.

[0006] The technical solution adopted in this invention is a wavelet Mamba low-light image enhancement method based on illumination prior, and the specific steps are as follows:

[0007] Step 1: Construct an illumination estimation module, taking the low-light image as input, to obtain the illumination quaternion prior. ;

[0008] Step 2: Construct a wavelet Mamba encoding / decoding network to integrate the low-light image with the illumination quaternion prior. The image is input into a wavelet Mamba encoding / decoding network to obtain the enhanced image;

[0009] Step 3: Train the overall network consisting of the illumination estimation module and the wavelet Mamba encoding / decoding network to obtain the trained overall network;

[0010] Step 4: Input the low-light image to be enhanced into the trained overall network to obtain the final enhanced image.

[0011] The invention is further characterized by:

[0012] In step 1, the specific processing procedure of the illumination estimation module is as follows:

[0013] Step 1.1: The input low-light image is processed through a learnable linear mapping layer. Shallow features are obtained, and the input low-light image is sequentially passed through a 1×1 convolution and a Gaussian filter. Processing yields enhanced features;

[0014] Step 1.2: Calculate the spectral intensity of the shallow features obtained in Step 1.1. Spectral slope and spectral curvature ;

[0015] Step 1.3, take the spectral intensity obtained in Step 1.2. Spectral slope and spectral curvature The weighted spectral intensity is obtained by multiplying each element-wise with the enhanced features obtained in step 1.1. Spectral slope and spectral curvature ;

[0016] Step 1.4, Calculate the hue prior. Priority of access Texture Prior Color Priority ;

[0017] Step 1.5, use the hue prior obtained in Step 1.4 Priority of access Texture Prior and color prior By splicing along the channel dimension, a quaternion of illumination priors is formed.

[0018] In step 1.2, the spectral intensity of shallow features The expression is:

[0019] (1)

[0020] In the formula, The spectrum of the light source, It is a specular reflection. For the material's reflectivity, The pixel spatial location represents the shallow layer features. ,in, x Represents the horizontal coordinate. y Represents the vertical coordinates;

[0021] Assuming equal-energy lighting, the spectrum of the light source Simplified to wavelength irrelevant Then we have:

[0022] (2)

[0023] Spectral slope of shallow features Spectral curvature of shallow features The expressions are as follows:

[0024] (3)

[0025] (4)

[0026] In equations (3) and (4), Indicates the reflectivity of a material The first-order spectral derivative of Gaussian filtering; Indicates the reflectivity of a material The second-order spectral derivative of Gaussian filtering.

[0027] In step 1.4, hue prior. The expression is:

[0028] (5);

[0029] In the formula, and All are learnable parameters;

[0030] Channel Prior The expression is:

[0031] (6);

[0032] The texture prior W is expressed as:

[0033] (7)

[0034] (8);

[0035] In the formula, For material reflectivity The Laplace operator, Set to 0.01 to prevent division by zero;

[0036] Color Priority The expression is:

[0037] (9)

[0038] In the formula, The normalized red channel component, The normalized green channel component, The normalized blue channel component;

[0039] , , By using the spectral intensity obtained in step 1.3 The RGB color channels are linearly normalized respectively. The pixel value at that location is mapped to the interval [-1, 1].

[0040] In step 2, the wavelet Mamba encoding and decoding network includes a first convolutional layer, an encoder, an illumination prior fusion module, a decoder, and an output convolutional layer;

[0041] The specific processing procedure of the wavelet Mamba encoding / decoding network is as follows:

[0042] Step 2.1: The H×W×C low-light image is passed through the first convolutional layer to extract initial features. These initial features are then input into the encoder to obtain... Feature map;

[0043] Step 2.2, take the result obtained in step 2.1 Feature maps and illumination quaternion priors The input is fed into the illumination prior fusion module to obtain illumination prior fusion features. ;

[0044] Step 2.3: Fuse the illumination prior features obtained in Step 2.2. The output of the third wavelet Mamba module in step 2.1 is sequentially input into the decoder and the output convolutional layer for processing to obtain the enhanced image.

[0045] In step 2.1, the kernel of the first convolutional layer is 3×3;

[0046] The encoder consists of a first wavelet Mamba module, a first downsampling module, a second wavelet Mamba module, a second downsampling module, a third wavelet Mamba module, and a third downsampling module in sequence.

[0047] The processing procedure for each wavelet Mamba module is as follows:

[0048] The input features are normalized and then linearly transformed using the data_transform function to map them to the interval [-1, 1] to obtain the mapped features. The mapped features are then decomposed into low-frequency and high-frequency components using Haar wavelet transform.

[0049] The low-frequency components are sequentially subjected to 3×3 convolution and ReLU activation to obtain the first processed features. The first processed features are then sequentially passed through a second convolutional layer, a ReLU activation function, and a third convolutional layer to obtain the first convolutional features. The first convolutional features are then reshaped into [batch_size, sequence_length, d_model] in a channel-first manner. Here, the reshaped sequence_length is obtained by reshaping the number of channels, and the reshaped d_model is the product of the height and width before reshaping. The reshaped features are then input into the first state space model for processing to obtain the first updated features. The state space model uses the mamba_ssm library from the third-party library Mamba and scans along the sequence length. The first updated features are then sequentially subjected to 3×3 convolution and layer normalization to obtain the enhanced low-frequency components.

[0050] After performing 3×3 convolution and ReLU activation operations on the high-frequency components sequentially, the second processed features are obtained. The second processed features are then passed through the fourth convolutional layer, the ReLU activation function, and the fifth convolutional layer to obtain the second convolutional features. The second convolutional features are reshaped into [batch_size, sequence_length, d_model] in a spatial dimension-first manner, where the reshaped sequence_length is the product of the height and width, and the reshaped d_model is obtained by reshaping the number of channels. The reshaped features are then input into the second state space model for processing to obtain the second updated features. The second state space model uses the mamba_ssm library from the third-party library Mamba. The second updated features are then subjected to 3×3 convolution and layer normalization operations sequentially to obtain the enhanced high-frequency components.

[0051] The enhanced low-frequency and high-frequency components are subjected to inverse discrete wavelet transform to obtain the hidden features of frequency enhancement. .

[0052] In step 2.2, the processing procedure of the illumination prior fusion module is as follows:

[0053] The result obtained in step 2.1 The feature maps are sequentially normalized and Fourier transformed to obtain the Query vector, which is then used to perform the illumination quaternion prior obtained in step 1. Perform layer normalization and Fourier transform operations sequentially to obtain the Key vector, and then use the illumination quaternion prior obtained in step 1. The layer normalization and convolution operations are performed sequentially to obtain the Value value. Latent space features are then obtained using the Query vector, Key vector, and Value value. These latent space features are then linearly transformed and... The feature maps are added together to obtain the illumination prior fused features. ;

[0054] The expressions for the query vector, key vector, value, and latent space features are as follows:

[0055] (5)

[0056] In the formula, Q represents the Query vector; K represents the Key vector; V represents the Value; FFT represents the Fourier transform; and SoftMax represents the SoftMax activation function. , and These are learnable, unbiased linear projection parameters.

[0057] In step 2.3, the decoder consists of a first upsampling module, a fourth wavelet Mamba module, a second upsampling module, a fifth wavelet Mamba module, a third upsampling module, and a sixth wavelet Mamba module in sequence.

[0058] The kernel of the output convolutional layer is 3×3;

[0059] The specific process is as follows:

[0060] Illumination prior fusion features The first upsampling input is used as the input. The output of the third wavelet Mamba module is added to the output of the first upsampling and used as the input of the fourth wavelet Mamba module. The output of the fourth wavelet Mamba module is used as the input of the second upsampling. The output of the second wavelet Mamba module is added to the output of the second upsampling and used as the input of the fifth wavelet Mamba module. The output of the fifth wavelet Mamba module is used as the input of the third upsampling. The output of the first wavelet Mamba module is added to the output of the third upsampling and used as the input of the sixth wavelet Mamba module. The output of the sixth wavelet Mamba module is fed into the output convolutional layer to obtain the enhanced image.

[0061] The total loss function during training is:

[0062] (10)

[0063] In the formula, These are hyperparameters, i = 1, 2, 3, 4; This is the L-1 loss value; For reconstruction losses; To perceive loss; To smooth out the loss;

[0064] in,

[0065] (11)

[0066] In equation (11), It is the low-frequency component of the input image after Haar wavelet transform decomposition. These are the low-frequency components of the real data;

[0067] (12)

[0068] In the formula, For image space pixel coordinates, To enhance the image, This is a true value image;

[0069] (13)

[0070] In the formula, This represents a VGG network pre-trained using the ImageNet dataset; Features of the ground truth image extracted by the VGG network. Features of the enhanced image extracted by the VGG network;

[0071] (14)

[0072] In the formula, and They represent in and gradient of direction, This indicates an enhanced image.

[0073] The beneficial effects of this invention are:

[0074] (1) The wavelet Mamba low-light image enhancement method based on illumination prior of the present invention provides multi-dimensional image prior through illumination estimation module, which effectively improves illumination distribution and enhances detail recovery capability;

[0075] (2) The present invention is a wavelet Mamba low-light image enhancement method based on illumination prior. The wavelet Mamba module is used to perform channel sensing and spatial sensing processing on low-frequency and high-frequency information in the frequency domain, respectively, to balance sensitive areas and solve long-range dependence, while achieving lightweighting in terms of parameter quantity and computational quantity.

[0076] (3) The wavelet Mamba low-light image enhancement method based on illumination prior of the present invention utilizes the illumination quaternary prior information through the illumination prior fusion module to enhance the image while maintaining naturalness and detail, resulting in high image quality after enhancement. Attached Figure Description

[0077] Figure 1 This is a flowchart of the wavelet Mamba low-light image enhancement method based on illumination prior of the present invention;

[0078] Figure 2 This is a schematic diagram of the illumination estimation module in the wavelet Mamba low-light image enhancement method based on illumination prior of the present invention;

[0079] Figure 3 This is a schematic diagram of the wavelet Mamba module in the wavelet Mamba low-light image enhancement method based on illumination prior of the present invention;

[0080] Figure 4 This is a schematic diagram of the illumination prior fusion module in the wavelet Mamba low-light image enhancement method based on illumination prior of the present invention;

[0081] Figure 5 The results show the comparison between the method of this invention and existing low-light image enhancement methods on the LOL-v2-real dataset;

[0082] Figure 6 The results show the experimental generalization performance of the method of this invention and existing low-light image enhancement methods on the IME and DICM datasets. Detailed Implementation

[0083] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0084] Example 1

[0085] This invention relates to a wavelet Mamba low-light image enhancement method based on illumination prior, such as... Figure 1 As shown, the specific steps are as follows:

[0086] Step 1: Construct an illumination estimation module, taking the low-light image as input, to obtain the illumination quaternion prior. ;

[0087] Step 2: Construct a wavelet Mamba encoding / decoding network to integrate the low-light image with the illumination quaternion prior. The image is input into a wavelet Mamba encoding / decoding network to obtain the enhanced image;

[0088] Step 3: Train the overall network consisting of the illumination estimation module and the wavelet Mamba encoding / decoding network to obtain the trained overall network;

[0089] Step 4: Input the low-light image to be enhanced into the trained overall network to obtain the final enhanced image.

[0090] Example 2

[0091] Based on Example 1, in step 1, the illumination estimation module includes a learnable linear mapping layer. 1×1 convolution, Gaussian filter Element-wise multiplication and concatenation operations;

[0092] like Figure 2 As shown, the specific processing procedure of the illumination estimation module is as follows:

[0093] Step 1.1: The input low-light image is processed through a learnable linear mapping layer. Shallow features are obtained, and the input low-light image is sequentially passed through a 1×1 convolution and a Gaussian filter. Processing yields enhanced features;

[0094] in, Defined using the torch.nn.Conv2d function, specifying 3 input channels, 3 output channels, and a 1×1 kernel size;

[0095] Gaussian filter Initialize using the np.exp function from the NumPy library, and then filter and enhance the features after the 1×1 convolution.

[0096] Step 1.2: Calculate the spectral intensity of the shallow features obtained in Step 1.1. Spectral slope and spectral curvature ;

[0097] Spectral intensity of shallow features The expression is:

[0098] (1)

[0099] In the formula, The spectrum of the light source, It is a specular reflection. For the material's reflectivity, The pixel spatial location represents the shallow layer features. ,in, x Represents the horizontal coordinate. y Represents the vertical coordinates;

[0100] Assuming equal-energy lighting, the spectrum of the light source Simplified to wavelength irrelevant Then we have:

[0101] (2)

[0102] Spectral slope of shallow features Spectral curvature of shallow features The expressions are as follows:

[0103] (3)

[0104] (4)

[0105] In equations (3) and (4), Indicates the reflectivity of a material The first-order spectral derivative of Gaussian filtering; Indicates the reflectivity of a material The second-order spectral derivative of Gaussian filtering;

[0106] Step 1.3, take the spectral intensity obtained in Step 1.2. Spectral slope and spectral curvature The weighted spectral intensity is obtained by multiplying each element-wise with the enhanced features obtained in step 1.1. Spectral slope and spectral curvature ;

[0107] Step 1.4, Calculate the hue prior. Priority of access Texture Prior Color Priority ;

[0108] Among them, hue a priori The expression is:

[0109] (5);

[0110] In the formula, and All are learnable parameters;

[0111] Channel Prior The expression is:

[0112] (6);

[0113] The texture prior W is expressed as:

[0114] (7)

[0115] (8);

[0116] In the formula, For material reflectivity The Laplace operator; Set to 0.01 to prevent division by zero;

[0117] in, In the formula, r Pixel spatial location representing shallow features ,wavelength The intensity of reflected light at that location; i Indicates the intensity of the incident light;

[0118] Color Priority The expression is:

[0119] (9)

[0120] In the formula, The normalized red channel component, The normalized green channel component, The normalized blue channel component;

[0121] , , By using the spectral intensity obtained in step 1.3 The RGB color channels are linearly normalized respectively. The pixel value at that location is mapped to the interval [-1, 1] to obtain the result;

[0122] Step 1.5, use the hue prior obtained in Step 1.4 Priority of access Texture Prior and color prior By splicing along the channel dimension, a quaternary prior for illumination is formed. .

[0123] Example 3

[0124] Based on Example 2, in step 2, the wavelet Mamba encoding and decoding network includes a first convolutional layer, an encoder, an illumination prior fusion module, a decoder, and an output convolutional layer;

[0125] The specific processing procedure of the wavelet Mamba encoding / decoding network is as follows:

[0126] Step 2.1: The H×W×C low-light image is passed through the first convolutional layer to extract initial features. These initial features are then input into the encoder to obtain... Feature map;

[0127] The kernel of the first convolutional layer is 3×3;

[0128] The encoder consists of a first wavelet Mamba module, a first downsampling module, a second wavelet Mamba module, a second downsampling module, a third wavelet Mamba module, and a third downsampling module in sequence.

[0129] like Figure 3 As shown, the processing procedure for each wavelet Mamba module is as follows:

[0130] The input features are normalized and then linearly transformed using the data_transform function to map them to the interval [-1, 1] to obtain the mapped features. The mapped features are then decomposed into low-frequency and high-frequency components using Haar wavelet transform.

[0131] The low-frequency components are sequentially subjected to 3×3 convolution and ReLU activation to obtain the first processed features. The first processed features are then sequentially passed through a second convolutional layer, a ReLU activation function, and a third convolutional layer to obtain the first convolutional features. The first convolutional features are then reshaped in a channel-first manner into [batch_size, sequence_length, d_model], where the reshaped sequence_length is obtained by reshaping the channels before reshaping, and the reshaped d_model is the product of the height and width before reshaping. The reshaped features are then input into the first state space model for processing to obtain the first updated features. The state space model uses the mamba_ssm library from the third-party library Mamba and scans along the sequence length. The input and output feature dimensions of the first state space model are both equal to the number of channels of the input features. The state expansion factor of the first state space model is 32, the local convolution width is 4, and the expansion factor is 2. The first updated features are then sequentially subjected to 3×3 convolution and layer normalization operations to obtain the enhanced low-frequency components.

[0132] After performing 3×3 convolution and ReLU activation operations on the high-frequency components sequentially, the second processed features are obtained. These second processed features are then passed through a fourth convolutional layer, a ReLU activation function, and a fifth convolutional layer to obtain the second convolutional features. These second convolutional features are then reshaped in a spatial dimension-first manner to [batch_size, sequence_length, d_model], that is, the second convolutional features are reshaped from their original shape [batch_size, height, width, channels] to [batch_size, sequence_length, d_model]. To reshape the product of height and width, the reshaped d_model is obtained by reshaping the channels. The reshaped features are input into the second state space model for processing to obtain the second updated features. The second state space model uses mamba_ssm from the third-party library Mamba. The input and output feature dimensions of the second state space model are equal to the number of channels of the input features. The state expansion factor of the second state space model is 32, the local convolution width is 4, and the expansion factor is 2. The second updated features are then subjected to 3×3 convolution and layer normalization operations to obtain the enhanced high-frequency components.

[0133] The enhanced low-frequency and high-frequency components are subjected to inverse discrete wavelet transform to obtain the hidden features of frequency enhancement. ;

[0134] The expression is:

[0135] ;

[0136] The input features of the first wavelet Mamba module are: The output features of the first wavelet Mamba module are: The first downsampling output feature is The input features of the second wavelet Mamba module are The output features of the second wavelet Mamba module are: The second downsampling output feature is The input features of the third wavelet Mamba module are The output features of the third wavelet Mamba module are: The third downsampling output feature is ;

[0137] Step 2.2, take the result obtained in step 2.1 Feature maps and illumination quaternion priors The input is fed into the illumination prior fusion module to obtain illumination prior fusion features. ;

[0138] like Figure 4 As shown, the processing procedure of the illumination prior fusion module is as follows:

[0139] The result obtained in step 2.1 The feature maps are sequentially normalized and Fourier transformed to obtain the Query vector, which is then used to perform the illumination quaternion prior obtained in step 1. Perform layer normalization and Fourier transform operations sequentially to obtain the Key vector, and then use the illumination quaternion prior obtained in step 1. The layer normalization and convolution operations are performed sequentially to obtain the Value value. Latent space features are then obtained using the Query vector, Key vector, and Value value. These latent space features are then linearly transformed and... The feature maps are added together to obtain the illumination prior fused features. ;

[0140] The expressions for the query vector, key vector, value, and latent space features are as follows:

[0141] (5)

[0142] In the formula, Q represents the Query vector; K represents the Key vector; V represents the Value; FFT represents the Fourier transform; and SoftMax represents the SoftMax activation function. , and These are learnable, unbiased linear projection parameters;

[0143] The core code snippet for calculating the vector product of Query and Key is: torch.matmul(query_fft, key_fft.conj().T) / self.scale;

[0144] The core code snippet for the SoftMax activation function is: attention_weights = F.softmax(attention_scores, dim=-1);

[0145] The core code snippet for performing a vector product with Value is: attention_output_fft = torch.matmul(attention_weights, value_fft);

[0146] Step 2.3: Fuse the illumination prior features obtained in Step 2.2. The output of the third wavelet Mamba module in step 2.1 is sequentially input into the decoder and the output convolutional layer for processing to obtain the enhanced image;

[0147] The decoder consists of a first upsampling module, a fourth wavelet Mamba module, a second upsampling module, a fifth wavelet Mamba module, a third upsampling module, and a sixth wavelet Mamba module.

[0148] The specific process is as follows:

[0149] Illumination prior fusion features The first upsampling input is used as the input, the output of the third wavelet Mamba module is added to the output of the first upsampling and used as the input of the fourth wavelet Mamba module, the output of the fourth wavelet Mamba module is used as the input of the second upsampling, the output of the second wavelet Mamba module is added to the output of the second upsampling and used as the input of the fifth wavelet Mamba module, the output of the fifth wavelet Mamba module is used as the input of the third upsampling, the output of the first wavelet Mamba module is added to the output of the third upsampling and used as the input of the sixth wavelet Mamba module, and the output of the sixth wavelet Mamba module is input into the output convolutional layer to obtain the enhanced image;

[0150] The processing flow of the fourth, fifth, and sixth wavelet Mamba modules is the same as that of the wavelet Mamba module in the encoder.

[0151] The input to the fourth wavelet Mamba module is The input to the fifth wavelet Mamba module is The input to the sixth wavelet Mamba module is H×W×C;

[0152] The kernel of the output convolutional layer is 3×3.

[0153] Example 4

[0154] Based on Example 3, the specific process of step 3 is as follows: When training the network using the LOLv1 and LOL-v2-real datasets, the LOLv1 and LOL-v2-real datasets are divided into training and testing sets; the LOLv1 dataset is divided into 485 pairs of training data and 15 pairs of testing data; the LOL-v2-real subset includes 689 pairs of images used for training and testing; the Adam optimizer is selected as the network optimizer, and the momentum of the Adam optimizer is set to 0.9, with an initial learning rate of 0.5; a cosine annealing scheme is used to adjust the learning rate, gradually reducing it according to a preset periodic pattern, reaching a minimum of 0.0001; during the training process, the entire training dataset is iterated 500 times, and hyperparameters such as the learning rate and the number of training epochs are dynamically adjusted by observing the total loss function index. Set to 0.01, 0.1, 0.1, 0.01.

[0155] Example 5

[0156] Based on Example 4, in step 3, the total loss function used during training is:

[0157] (10)

[0158] In the formula, These are hyperparameters, i = 1, 2, 3, 4; This is the L-1 loss value; For reconstruction losses; To perceive loss; To smooth out the loss;

[0159] in,

[0160] (11)

[0161] In equation (11), It is the low-frequency component of the input image after Haar wavelet transform decomposition. These are the low-frequency components of the real data;

[0162] The core code snippet is: L_low = self.low_frequency_loss(F_low, G_low);

[0163] (12)

[0164] In the formula, For image space pixel coordinates, To enhance the image, This is a true value image;

[0165] The core code snippet is: L_rec = self.reconstruction_loss(I_low, I_gt);

[0166] (13)

[0167] In the formula, This represents a VGG network pre-trained using the ImageNet dataset; Features of the ground truth image extracted by the VGG network. Features of the enhanced image extracted by the VGG network;

[0168] The core code snippet is: L_per = self.perceptual_loss(VGG_features_low, VGG_features_gt);

[0169] (14)

[0170] In the formula, and They represent in and gradient of direction, Indicates an enhanced image;

[0171] By using the SummaryWriter function of the Python third-party library TensorBoard, key metrics such as the total loss during training are periodically output and recorded in the TensorBoard log file. At each stage of training, the network parameters, the current training epoch, and the optimizer state are saved to obtain the overall network weight file after training.

[0172] Example 6

[0173] Using Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and Learning-Based Perceptual Image Patch Similarity (LPIPS) metrics, the present invention and existing methods were compared and analyzed on the LOLv1 dataset for low-light enhancement, and the test results are shown in Table 1. The present invention and existing methods were also compared and analyzed on the LOL-v2-real dataset for low-light enhancement, and the test results are shown in Table 2. Tables 1 and 2 show that the method of the present invention has great potential in the field of low-light enhancement. The present invention outperforms other competing methods and can achieve excellent results on the datasets (e.g., Figure 5 (As shown).

[0174] Table 1 Comparison of metrics between the present invention and existing methods on the LOLv1 dataset.

[0175]

[0176] Table 2 Comparison of metrics between the present invention and existing methods on the LOL-v2-real dataset

[0177]

[0178] The present invention also compared the model parameter count, number of floating-point operations, and processing speed of the method of the present invention with those of existing low-light enhancement methods. As shown in Table 3, the present invention has a significant advantage in processing speed, while also performing well in terms of parameter count and computational load. It has excellent overall performance and is suitable for application scenarios that require fast image processing.

[0179] Table 3. Comparison of the method of the present invention with existing methods in terms of parameter quantity, computational load, and processing speed.

[0180]

[0181] Figure 6 The enhancement results are shown on the IME and DICM datasets (which have no reference images). It can be seen that the present invention can adapt to the restoration under different lighting conditions, and the visual effect is more natural and clearer than the LLFormer method.

Claims

1. A wavelet Mamba low-light image enhancement method based on illumination prior, characterized in that, The specific steps are as follows: Step 1: Construct an illumination estimation module, taking the low-light image as input, to obtain the illumination quaternion prior. ; In step 1, the specific processing procedure of the illumination estimation module is as follows: Step 1.1: The input low-light image is processed through a learnable linear mapping layer. Shallow features are obtained, and the input low-light image is sequentially passed through a 1×1 convolution and a Gaussian filter. Processing yields enhanced features; Step 1.2: Calculate the spectral intensity of the shallow features obtained in Step 1.

1. Spectral slope and spectral curvature ; in, The pixel spatial location represents the shallow layer features. , x Represents the horizontal coordinate. y Represents the vertical coordinates; Indicates wavelength; Step 1.3, take the spectral intensity obtained in Step 1.

2. Spectral slope and spectral curvature The weighted spectral intensity is obtained by multiplying each element-wise with the enhanced features obtained in step 1.

1. Spectral slope and spectral curvature ; Step 1.4, Calculate the hue prior. Priority of access Texture Prior Color Priority ; In step 1.4, hue prior. The expression is: (5); In the formula, and All are learnable parameters; Channel Prior The expression is: (6); In the formula, Indicates the reflectivity of a material The first-order spectral derivative of Gaussian filtering; Indicates the reflectivity of a material The second-order spectral derivative of Gaussian filtering; The texture prior W is expressed as: (7) (8); In the formula, For material reflectivity The Laplace operator, Set to 0.01 to prevent division by zero; Color Priority The expression is: (9) In the formula, The normalized red channel component, The normalized green channel component, The normalized blue channel component; , , By using the spectral intensity obtained in step 1.3 The RGB color channels are linearly normalized respectively. The pixel value at that location is mapped to the interval [-1, 1] to obtain the result; Step 1.5, use the hue prior obtained in Step 1.4 Priority of access Texture Prior and color prior The light quaternion priors are spliced ​​together along the channel dimension; Step 2: Construct a wavelet Mamba encoding / decoding network to integrate the low-light image with the illumination quaternion prior. The image is input into a wavelet Mamba encoding / decoding network to obtain the enhanced image; In step 2, the wavelet Mamba encoding and decoding network includes a first convolutional layer, an encoder, an illumination prior fusion module, a decoder, and an output convolutional layer; The specific processing procedure of the wavelet Mamba encoding / decoding network is as follows: Step 2.1: The H×W×C low-light image is passed through the first convolutional layer to extract initial features. These initial features are then input into the encoder to obtain... Feature map; Step 2.2, take the result obtained in step 2.1 Feature maps and illumination quaternion priors The input is fed into the illumination prior fusion module to obtain illumination prior fusion features. ; The processing procedure of the illumination prior fusion module is as follows: The result obtained in step 2.1 The feature maps are sequentially normalized and Fourier transformed to obtain the Query vector, which is then used to perform the illumination quaternion prior obtained in step 1. Perform layer normalization and Fourier transform operations sequentially to obtain the Key vector, and then use the illumination quaternion prior obtained in step 1. The layer normalization and convolution operations are performed sequentially to obtain the Value value. Latent space features are then obtained using the Query vector, Key vector, and Value value. These latent space features are then linearly transformed and... The feature maps are added together to obtain the illumination prior fused features. ; Step 2.3: Fuse the illumination prior features obtained in Step 2.

2. The output of the third wavelet Mamba module in step 2.1 is sequentially input into the decoder and the output convolutional layer for processing to obtain the enhanced image; Step 3: Train the overall network consisting of the illumination estimation module and the wavelet Mamba encoding / decoding network to obtain the trained overall network; Step 4: Input the low-light image to be enhanced into the trained overall network to obtain the final enhanced image.

2. The wavelet Mamba low-light image enhancement method based on illumination prior as described in claim 1, characterized in that, In step 1.2, the spectral intensity of shallow features The expression is: (1) In the formula, The spectrum of the light source, It is a specular reflection. The material's reflectivity; Assuming equal-energy lighting, the spectrum of the light source Simplified to wavelength irrelevant Then we have: (2) Spectral slope of shallow features Spectral curvature of shallow features The expressions are as follows: (3) (4)。 3. The wavelet Mamba low-light image enhancement method based on illumination prior as described in claim 1, characterized in that, In step 2.1, the kernel of the first convolutional layer is 3×3; The encoder consists of a first wavelet Mamba module, a first downsampling module, a second wavelet Mamba module, a second downsampling module, a third wavelet Mamba module, and a third downsampling module in sequence. The processing procedure for each wavelet Mamba module is as follows: The input features are normalized and then linearly transformed using the data_transform function to map them to the interval [-1, 1] to obtain the mapped features. The mapped features are then decomposed into low-frequency and high-frequency components using Haar wavelet transform. The low-frequency components are sequentially subjected to 3×3 convolution and ReLU activation to obtain the first processed features. The first processed features are then sequentially passed through a second convolutional layer, a ReLU activation function, and a third convolutional layer to obtain the first convolutional features. The first convolutional features are then reshaped into [batch_size, sequence_length, d_model] in a channel-first manner. Here, the reshaped sequence_length is obtained by reshaping the number of channels, and the reshaped d_model is the product of the height and width before reshaping. The reshaped features are then input into the first state space model for processing to obtain the first updated features. The state space model uses the mamba_ssm library from the third-party library Mamba and scans along the sequence length. The first updated features are then sequentially subjected to 3×3 convolution and layer normalization to obtain the enhanced low-frequency components. After performing 3×3 convolution and ReLU activation operations on the high-frequency components sequentially, the second processed features are obtained. The second processed features are then passed through the fourth convolutional layer, the ReLU activation function, and the fifth convolutional layer to obtain the second convolutional features. The second convolutional features are reshaped into [batch_size, sequence_length, d_model] in a spatial dimension-first manner, where the reshaped sequence_length is the product of the original height and width, and the reshaped d_model is obtained by reshaping the number of channels. The reshaped features are then input into the second state space model for processing to obtain the second updated features. The second state space model uses the mamba_ssm library from the third-party library Mamba. The second updated features are then subjected to 3×3 convolution and layer normalization operations sequentially to obtain the enhanced high-frequency components. The enhanced low-frequency and high-frequency components are subjected to inverse discrete wavelet transform to obtain the hidden features of frequency enhancement. .

4. The wavelet Mamba low-light image enhancement method based on illumination prior as described in claim 1, characterized in that, In step 2.2, The expressions for the query vector, key vector, value, and latent space features are as follows: (5) In the formula, Q represents the Query vector; K represents the Key vector; V represents the Value; FFT represents the Fourier transform; and SoftMax represents the SoftMax activation function. , and These are learnable, unbiased linear projection parameters.

5. The wavelet Mamba low-light image enhancement method based on illumination prior as described in claim 1, characterized in that, In step 2.3, the decoder consists of a first upsampling module, a fourth wavelet Mamba module, a second upsampling module, a fifth wavelet Mamba module, a third upsampling module, and a sixth wavelet Mamba module in sequence. The kernel of the output convolutional layer is 3×3; The specific process is as follows: Illumination prior fusion features The first upsampling input is used as the input. The output of the third wavelet Mamba module is added to the output of the first upsampling and used as the input of the fourth wavelet Mamba module. The output of the fourth wavelet Mamba module is used as the input of the second upsampling. The output of the second wavelet Mamba module is added to the output of the second upsampling and used as the input of the fifth wavelet Mamba module. The output of the fifth wavelet Mamba module is used as the input of the third upsampling. The output of the first wavelet Mamba module is added to the output of the third upsampling and used as the input of the sixth wavelet Mamba module. The output of the sixth wavelet Mamba module is fed into the output convolutional layer to obtain the enhanced image.

6. The wavelet Mamba low-light image enhancement method based on illumination prior as described in claim 1, characterized in that, The total loss function during training is: (10) In the formula, These are hyperparameters, i = 1, 2, 3, 4; This is the L-1 loss value; For reconstruction losses; To perceive loss; To smooth out the loss; in, (11) In equation (11), It is the low-frequency component of the input image after Haar wavelet transform decomposition. These are the low-frequency components of the real data; (12) In the formula, For image space pixel coordinates, To enhance the image, This is a true value image; (13) In the formula, This represents a VGG network pre-trained using the ImageNet dataset; Features of the ground truth image extracted by the VGG network. Features of the enhanced image extracted by the VGG network; (14) In the formula, and They represent in and gradient of direction, This indicates an enhanced image.

Citation Information

Patent Citations

  • Unsupervised low-illumination image enhancement method and system based on physical prior

    CN118229596A

  • Low-illumination image enhancement method and system based on optical prior and spectrum constraint

    CN119762368A