Transform illumination migration low-light image enhancement method for zero sample based on wavelet transform guidance

Through wavelet transformation and Retinex theory combined with the zero-sample light migration method of Transformer network, the problem of insufficient generalization ability of low-light images under complex lighting conditions is solved, efficient image enhancement effect is achieved, and image quality and adaptability are improved.

CN120374428APending Publication Date: 2025-07-25CHONGQING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510503641.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing low-light image enhancement methods have limited generalization capabilities under complex lighting conditions, making it difficult to effectively improve image quality and adapt to different lighting environments, especially in zero-sample learning scenarios.

Method used

The Transformer illumination migration method based on wavelet transformation guided by zero-sample wavelet transformation is used to extract the illumination changes in the low-frequency domain through discrete wavelet transformation. Combined with Retinex theory and Transformer network, the image is decomposed into reflectivity and lighting diagrams, and the self-attention module and the light migration network are used for feature recombination and enhancement.

Benefits of technology

It significantly improves the visual quality and generalization ability of low-light images, can achieve robust enhancement effects under unknown lighting conditions, maintain image content structure and reduce noise interference, and improves the interpretability and robustness of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374428A_ABST
    Figure CN120374428A_ABST
Patent Text Reader

Abstract

The invention discloses a zero-sample Transform illumination migration low-light image enhancement method based on wavelet transform guidance, and relates to the technical field of low-light image processing. According to the zero-sample low-light image enhancement method based on the Transform, the illumination prior of the low-frequency domain is obtained by exploring the illumination change of the wavelet transform on the low-frequency domain, and the robust and effective enhancement performance is realized in combination with the Transform illumination migration network; a preliminary illumination enhancement network is designed based on a Retinex theory to obtain a preliminary enhanced image, so that the acquisition of illumination prior in a subsequent illumination migration process is enhanced, the content structure of a low-light image is further cleared, and the interference of disordered data is reduced; wide experiments are carried out based on a real world data set, and the effectiveness and generalization of the method are proved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of low-light image processing, and specifically to a zero-sample wavelet transform-guided Transformer illumination transfer low-light image enhancement method. Background Art

[0002] Images captured in low-light environments are usually affected by degradation factors such as reduced visibility and noise interference, which seriously weaken the image quality and have an adverse impact on the performance of downstream computer vision tasks such as autonomous driving, object detection, and image classification. To address this issue, researchers have proposed various low-light image enhancement (LLIE) techniques, aiming to restore the detailed information in under-illuminated areas and improve the overall image quality to effectively enhance the robustness of subsequent vision tasks. Traditional methods, such as histogram equalization (HE), Retinex theory, and gamma correction, mainly rely on the internal characteristics of the image for optimization. Although the image quality has been improved to a certain extent, due to these methods being based on manually designed prior knowledge and lacking adaptability, it is difficult to cope with complex and variable lighting conditions, resulting in instability and performance bottlenecks in the enhancement effect.

[0003] As an important and challenging task in the field of computer vision, low-light image enhancement (LLIE) has received extensive attention in recent years. With the rise of deep learning, traditional methods have gradually been replaced by deep learning-based technologies. However, the mainstream methods in current research still mainly rely on supervised learning, relying on a large amount of paired data for training to model the lighting changes and environmental factors in the real world, thereby improving the enhancement effect. For example, Cai et al. proposed a single-stage framework based on the Retinex theory and incorporated Transformer to model non-local interaction relationships, thereby enhancing the enhancement performance. He et al. first introduced visual language prompts in the wavelet diffusion model for iterative guidance to strengthen the constraint on the enhancement result. Although these methods have made significant progress, their high dependence on paired data limits their generalization ability under unforeseen lighting conditions, affecting their robustness.

[0004] Different from linear image restoration tasks, low-light image enhancement faces more complex unknown degradation factors, such as low signal-to-noise ratio and uneven illumination, making it extremely challenging to directly restore clean images. The core objective of low-light image enhancement is to learn the mapping relationship from the degradation domain to the normal domain to eliminate the impact of poor illumination and improve image quality. Currently, deep learning-based methods have become the mainstream, which rely on data-driven ways for feature learning and have improved the enhancement effect to varying degrees. For example, RetinexNet based on the Retinex theory combines with a deep neural network to enhance image details through illumination-reflection decomposition, while KinD further uses a multi-stage optimization strategy to improve visual quality. With the construction of large-scale low-light image datasets, more and more supervised methods have been proposed. For instance, DiffLL combines a diffusion model and wavelet transform to effectively reduce the challenges of high-cost inference of the diffusion model. GSAD proposes a global regularization method to constrain the diffusion model. However, these methods highly rely on paired training data, and their generalization ability under complex illumination conditions is still limited.

[0005] To reduce the dependence on paired data, researchers have started to explore unsupervised or unpaired learning methods. For example, EnlightGAN adopts a generative adversarial network (GANs) as the core framework to learn the distribution characteristics of low-light images through unpaired data, thus generating naturally enhanced images under different lighting environments. In addition, zero-shot learning methods have received extensive attention in recent years. For example, Zero-DCE adjusts the lighting distribution through a learnable curve, while SCI adopts a self-calibrated illumination optimization framework to achieve efficient and unsupervised low-light image enhancement. GDP and FourierDiff utilize the generative generalization ability of pre-trained diffusion models to strengthen the parameter bridging of the model. Although existing methods have improved the quality of low-light images to varying degrees, there are still challenges in balancing perceptual quality, detail preservation, and inference performance.

[0006] Based on the above problems, recent research has begun to explore how to reduce the dependence on supervision signals to improve the adaptability of low-light image enhancement methods in complex environments. Most previous works optimize image parameters through generative models and deep neural networks based on the Retinex theory. For example, EnlightenGAN and NeRCo adopt generative adversarial networks, use adversarial learning strategies for unsupervised training, and construct discriminators to guide the low-light enhancement model. LightingDiffusion combines the Retinex theory to perform illumination restoration at the feature level and refines it through a diffusion model, achieving effective results. However, although these unsupervised methods have reduced the dependence on supervision signals to a certain extent, they still retain the requirement for specific normal light training data, limiting their ability to be generalized to unknown scenarios.

[0007] Therefore, in order to further get rid of the dependence on supervision signals and normal light training data, a more direct zero-shot low-light image enhancement method has been gradually taken seriously. By seeking to utilize the illumination knowledge contained in the image itself for fitting. For example, ZeroDCE learns the image curve based on the Retinex theory to achieve high-efficiency enhancement. GDP and FourierDiff utilize the generation generalization ability of the pre-trained diffusion model to strengthen the parameter bridging of the model. However, most of these methods rely on a single Retinex network or generative model, resulting in large differences in performance or serious inference time costs.

[0008] In recent years, the wide application of Transformer in the field of computer vision has brought significant progress to the low-light image enhancement (LLIE) task. Its self-attention mechanism has natural advantages in capturing long-range dependencies and modeling non-local features. Compared with traditional convolutional neural networks (CNNs), it can more effectively learn complex illumination change patterns, thereby improving the enhancement effect. Existing research mainly focuses on two aspects: one is to combine the Retinex theory and use the self-attention mechanism to model the illumination and reflection components; the other is to construct an end-to-end Transformer enhancement framework to directly learn global features and achieve more accurate image enhancement. For example, RUAS models the illumination information through an adaptive attention module, thereby effectively enhancing underexposed images; LLFormer combines CNN and Swin Transformer and uses a hierarchical attention mechanism to model the illumination information at different scales to improve the enhancement quality; Restormer optimizes the feature aggregation strategy based on the Transformer structure to improve the image restoration effect; RetinexFormer proposes a single-stage framework based on the Retinex theory and integrates Transformer to model non-local interaction relationships to further improve the enhancement performance. However, existing Transformer-based frameworks still face major challenges in the zero-shot learning scenario. Previous methods have limitations in adapting to different illumination environments and are difficult to meet the enhancement requirements of high-quality visual perception. Therefore, how to improve the generalization ability of Transformer in unsupervised and zero-shot learning remains an urgent problem to be solved in the field of low-light image enhancement.

[0009] Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention

[0010] The purpose of the present invention is to provide a zero-shot wavelet transform-guided Transformer illumination transfer low-light image enhancement method to solve the technical problems proposed in the background art.

[0011] To achieve the above object, the present invention provides the following technical solutions: A zero-shot Transformer illumination transfer low-light image enhancement method guided by wavelet transform, which at least includes the following steps:

[0012] S1: The input low-light image is subjected to K discrete wavelet transforms (DWT) to obtain its low-frequency domain, which is jointly fed into an initial illumination enhancement network to separately extract the reflection map and the illumination map, and the initial enhanced image is restored.

[0013] S2: Subsequently, Perform K discrete wavelet transforms again, extract its low-frequency domain and resize it to obtain a high-brightness illumination image. After patch partitioning, both are jointly fed into the Transform Block for feature encoding, and the content structure f C and the illumination background f L are separately extracted, and then effective feature recombination and fusion are performed with the help of the Transform Light-Transfer Block;

[0014] S3: Finally, the final output image is restored via the decoder;

[0015] S4: In order to further constrain the illumination intensity and effect of the enhancement result, the perceptual effect of the illumination transfer process is measured by a loss function.

[0016] Furthermore, the initial illumination enhancement network is constructed based on the Retinex theory. The initial illumination enhancement network assumes that the input image I can be decomposed into a reflectance map R and an illumination map L. Among them, R represents the inherent color and texture information of the object, and L reflects the illumination conditions of the image, that is:

[0017] I = R·L

[0018] Therefore, the reflection map R and the illumination map L are separately extracted through two branches. The branch for extracting L takes the low-frequency domain of the input image after K discrete wavelet transforms and resizes it as the input. The subsequent operation methods are roughly the same in the two branches.

[0019] Furthermore, obtaining the initially restored enhanced image at least includes the following steps:

[0020] First, use two sets of blocks combined with several convolutions to obtain the embedded features and process them with the Relu activation function to obtain the reflection feature R1 and the illumination feature L1;

[0021] Subsequently, use the self-attention module to refine and further extract the content and illumination information in each branch to obtain the reflection feature R2 and the illumination feature L2;

[0022] Among them, in the reflection branch, illumination features are used to supplement the content information in the reflection features, while in the illumination branch, a residual structure is used to process redundant information to obtain clean illumination features;

[0023] Finally, the processing results of each branch are processed again by a convolution block, batch normalization, and Sigmoid activation function to obtain the final reflectance map R and illumination map L, expressed as:

[0024] R = Sig(BN(Convs(R2 + L2)))

[0025] L = Sig(BN(Convs(L1 - L2)))

[0026] Among them, Sig represents the Sigmoid activation function; BN represents batch normalization; Convs represents a convolution block composed of multiple convolution operations;

[0027] Subsequently, based on the obtained R and L, element-wise multiplication is performed in combination with the Retinex theory to obtain a preliminarily restored enhanced image

[0028] Furthermore, the S2 at least includes the following steps:

[0029] First, Perform K discrete wavelet transforms again, extract its low-frequency domain and resize it to obtain a high-brightness illumination image

[0030] Subsequently, As the content image, As the illumination image, after patch partitioning, they are jointly fed into the Transform Block for feature encoding to extract content and illumination features respectively;

[0031] Among them, each of the said Transform Blocks is a basic box composed of operations such as basic LayerNorm, Window Attention, and multi-layer perceptron. Specifically:

[0032] First, the image is partitioned into various non-overlapping n×n patches as tokens in the Transformer, and n is set to 2;

[0033] Subsequently, the partitioned patch blocks are sent into the Transform Block for LayerNorm operation, and the window attention module is applied for calculation. Finally, another LayerNorm operation is performed and projection mapping is processed through the MLP layer. And after each operation, we integrate residual connections to retain more feature information, formulating the process as:

[0034]

[0035] where Q, K, and V are the query, key, and value matrices respectively; d is the matrix dimension; WA M×M represents multi-head self-attention using a window of shape M×M; B is a bias term; WA 2n×2n represents multi-head self-attention using a window of shape 2n×2n; LN is the layer normalization operation; I l-1 represents the output of the previous layer; MLP represents the multi-layer perceptron;

[0036] After passing through three Transform Blocks of the same module, content feature f and are respectively extracted from C and illumination feature f L ;

[0037] Subsequently, based on the extracted features, they are sent into the Transform Light-Transfer Block for content and illumination integration to learn a specific illumination level for effective enhancement.

[0038] Furthermore, the output of S3 is based on the output features with good light enhancement obtained by S1 - S2 in the previous operation. Following the settings of S1 - S2, a Decoder constructed by VGG is used to decode the enhanced features to obtain the final output I out .

[0039] Furthermore, S4 at least includes the following steps:

[0040]

[0041] where represents the feature map after transforming the input image through a certain layer of feature extraction module; u(·) and represent the mean and variance of the features respectively;

[0042] Meanwhile, following the previous work, a consistency loss is also adopted to further maintain the structure of the content image and the illumination features of the illumination image, formulated as:

[0043]

[0044] Among them, I CC represents the output image obtained by extracting illumination and content from and integrating them; I LL represents the output image obtained by extracting illumination and content from and integrating them.

[0045] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0046] 1. The present invention proposes a zero-shot low-light image enhancement method based on Transformer. By exploring the illumination changes in the low-frequency domain through wavelet transform, its illumination prior is obtained, and a robust and effective enhancement performance is achieved by combining the Transform illumination migration network;

[0047] 2. The present invention designs a preliminary illumination enhancement network based on the Retinex theory to obtain a preliminary enhanced image, strengthens the acquisition of illumination prior in the subsequent illumination migration process, further clarifies the content structure of the low-light image, and reduces the interference of chaotic data;

[0048] 3. The extensive experiments carried out by the present invention on a real-world dataset prove the effectiveness and generalization of the proposed method. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 is the overall framework flowchart of the present invention;

[0051] Figure 2 is a visual comparison diagram of the low-light enhancement method on the LOLV1 dataset (the first row) and the LOLV2 dataset (the second row) of the present invention;

[0052] Figure 3 is a schematic visual comparison diagram of the low-light enhancement method used by the present invention on an unpaired dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, not all, embodiments of the present invention.

[0054] The present invention deeply explores the impact of wavelet transform on image illumination changes and discovers that after multiple transformations, significant illumination enhancement characteristics will emerge in the low-frequency components. Based on this observation, the present invention combines the Retinex theory to construct a preliminary illumination enhancement network to extract and strengthen the illumination information of the image. The present invention designs an illumination migration network. Based on the Transform module, it uses the preliminarily enhanced image to extract the content structure, and at the same time extracts the illumination background in the low-frequency domain after its multiple wavelet transformations and performs migration to achieve more effective adaptive enhancement. Based on the above-mentioned proposed model architecture for low-light image enhancement, a large number of experimental results show that this method exhibits good interpretability, robustness, and efficiency in various low-light scenarios, significantly improving the visual quality of low-light images.

[0055] Specifically as follows:

[0056] Please refer to Figure 1 , the zero-shot wavelet transform-guided Transformer illumination migration low-light image enhancement method, at least includes the following steps:

[0057] S1: Obtain the low-frequency domain of the input low-light image by performing K discrete wavelet transforms (DWT) and jointly feed it into the initial illumination enhancement network to separately extract the reflectance map and the illumination map, and restore the initial enhanced image

[0058] S2: Subsequently, Perform K discrete wavelet transforms again, extract its low-frequency domain and resize it to obtain a high-brightness illumination image After patch partitioning, both are jointly fed into the Transform Block for feature encoding to separately extract the content structure f C and the illumination background f L , and then perform effective feature recombination and fusion with the help of the Transform Light-TransferBlock;

[0059] S3: Finally, restore the final output image through the decoder;

[0060] S4: In order to further constrain the illumination intensity and effect of the enhancement result, the perceptual effect of the illumination migration process is measured by a loss function.

[0061] The initial illumination enhancement network is constructed based on the Retinex theory. The initial illumination enhancement network assumes that the input image I can be decomposed into a reflectance map R and an illumination map L. Among them, R represents the inherent color and texture information of the object, and L reflects the illumination conditions of the image, that is:

[0062] I = R·L

[0063] Therefore, the reflection map R and the illumination map L are respectively extracted through two branches. The branch for extracting L takes the low-frequency domain of the input image after K discrete wavelet transforms and resizes it as the input. The subsequent operation methods are roughly the same in the two branches.

[0064] Obtain the preliminarily restored enhanced image It includes at least the following steps:

[0065] First, use two sets of blocks composed of several convolutions to obtain the embedded features and process them using the Relu activation function to obtain the reflection feature R1 and the illumination feature L1;

[0066] Subsequently, use the self-attention module to refine and further extract the content and illumination information in each branch to obtain the reflection feature R2 and the illumination feature L2;

[0067] Among them, in the reflection branch, the illumination feature is used to supplement the content information in the reflection feature, while for the illumination branch, the residual structure is used to process the redundant information to obtain the clean illumination feature;

[0068] Finally, the processing results of each branch are processed again by the convolution block, batch normalization, and Sigmoid activation function to obtain the final reflectance map R and illumination map L, expressed as:

[0069] R = Sig(BN(Convs(R2 + L2)))

[0070] L = Sig(BN(Convs(L1 - L2)))

[0071] Among them, Sig represents the Sigmoid activation function; BN represents batch normalization; Convs represents the convolution block composed of multiple convolution operations;

[0072] Subsequently, based on the obtained R and L, perform element-wise multiplication in combination with the Retinex theory to obtain the preliminarily restored enhanced image

[0073] S2 includes at least the following steps:

[0074] First Perform K discrete wavelet transforms again, extract its low-frequency domain and resize it to obtain the high-brightness illumination image

[0075] Subsequently As the content image, As the illumination image, after patch partitioning, they are jointly fed into the Transform Block for feature encoding to respectively extract the content and illumination features;

[0076] Among them, each Transform Block is a basic box composed of operations such as basic LayerNorm, WindowAttention, and multi-layer perceptron. Specifically:

[0077] First, the image is partitioned into various non-overlapping n×n patches as tokens in the Transformer, and n is set to 2.

[0078] Subsequently, the partitioned patch blocks are sent into the Transform Block for LayerNorm operation, and the window attention module is applied for calculation. Finally, LayerNorm operation is performed again and projection mapping is processed through the MLP layer. And after each operation, we integrate residual connections to retain more feature information, formulating its process as:

[0079]

[0080] Among them, Q, K, and V are query, key, and value matrices respectively; d is the matrix dimension; WA M×M represents multi-head self-attention using an M×M shaped window; B is a bias term; WA 2n×2n represents multi-head self-attention using a 2n×2n shaped window; LN is the layer normalization operation; I l-1 represents the output of the previous layer; MLP represents the multi-layer perceptron;

[0081] After passing through three layers of Transform Block of the same module, content feature f and are respectively extracted from C and illumination feature f L ;

[0082] Subsequently, based on the extracted features, they are sent into the Transform Light-Transfer Block for content and illumination integration to learn a specific illumination level for effective enhancement. Specifically:

[0083] In the implementation, a close similarity to the original Transformer decoder is maintained, but there are two key changes: a) The initial attention module of each Transformer decoder layer is MSA, while previously MHA (multi-head attention) was adopted; B) LayerNorm is before the attention module and MLP, rather than after them.

[0084] The output of S3 is based on the output features with good light enhancement obtained by S1 - S2 in the previous operation. Following the settings of S1 - S2, a Decoder constructed by VGG is used to decode the enhanced features to obtain the final output I out。

[0085] S4 includes at least the following steps:

[0086]

[0087] Among them, represents the feature map after transforming the input image through a certain layer of feature extraction module; u(·) and respectively represent the mean and variance of the features;

[0088] Meanwhile, following previous work, a consistency loss is also adopted to further preserve the structure of the content image and the illumination features of the illumination image, formulated as:

[0089]

[0090] Among them, I CC represents the output image obtained by integrating the extracted illumination and content from ; I LL represents the output image obtained by integrating the extracted illumination and content from

[0091] In summary:

[0092] In the present invention, first, the influence of wavelet transform on the illumination change in its low-frequency domain is verified. On this basis, the present invention obtains the motivation to explore the transform in the field of zero-shot low-light image enhancement. Specifically, in order to strengthen the acquisition of illumination prior, the present invention first performs multiple wavelet transforms on the low-light image to obtain its low-frequency domain to construct a high-brightness illumination image, and further combines the Reintex theory to construct a preliminary illumination enhancement network to obtain a preliminary enhanced image. Subsequently, the preliminarily enhanced image is subjected to multiple wavelet transforms again to obtain its low-frequency domain. At this time, the low-frequency domain has more significant structural information and brightness prior, and the two images are fed into the illumination migration network constructed based on the transform box to respectively extract the content structure and brightness background. By effectively fusing them, the present invention realizes effective adaptive enhancement. It should be noted that since the wavelet low-frequency domain preserves the main content structure of the original image, it can better meet the data distribution of the degraded image and reduce the interference of chaotic data. Sufficient experiments show that this framework has generalization in various scenarios.

[0093] Based on the above method, the following specific experiments are proposed:

[0094] Experimental settings

[0095] ​The present invention implements the framework of the present invention using Pytorch on a single NVIDIA GeForce RTX 3090 GPU with a batch size of 8. The present invention sets the total number of training to 6×104 iterations. The present invention uses the Adam optimizer with a learning rate of 10-4. The number of wavelets converts k to 3.

[0096] Benchmark dataset

[0097] To verify the effectiveness of the method of the present invention, the present invention trained and evaluated it on the LOLV1 dataset. The LOLV1 dataset contains 500 pairs of real low-light / normal-light images, of which 485 pairs are used for training and 15 pairs are used for testing. By adjusting the exposure time and ISO while keeping other camera parameters unchanged, the low-light images in this dataset were collected to ensure the authenticity and diversity of the data. In addition, to enhance the comprehensiveness of the experiment, the present invention extended multiple real-world benchmark datasets to more thoroughly evaluate the generalization and robustness of the proposed method. Specifically, the present invention adopted an extended paired dataset, LOLV2 real capture, and three unpaired datasets, MEF, Lime, and DICM. To maintain consistency, the present invention collectively refers to these three unpaired datasets as the "unpaired set" and evaluates them under the same experimental settings. It is important to emphasize that throughout the training and testing process, the present invention only uses the low-light images in the paired dataset as input without using the corresponding normal-light images.

[0098] Metrics

[0099] For the paired dataset, the present invention uses two distortion metrics (PSNR and SSIM) and two perceptual metrics, LPIPS and FID, to evaluate the network. For the unpaired dataset, the present invention uses NIQE and PI as evaluation metrics. A higher PSNR or SSIM means a more realistic recovery result, while LPIPS, FID, NIQE, or PI indicate higher-quality details, more attractive visual perception, and higher color fidelity.

[0100] Comparison with the state-of-the-art

[0101] To verify the effectiveness of the method proposed by the present invention, the present invention compared it with unsupervised learning methods in recent years. For example: Zero-DCE++; RUAS; Enlightengan; SCI; Clip-lit; Pairlie; GDP; Zeroled; Fourierdiff and LightEdendEddiffusion.

[0102] Quantitative comparison

[0103] The present invention runs its publicly available official code through a verified model, thereby obtaining the quantitative results of other methods. As shown in Table 1, the method of the present invention achieves performance comparable to the state of the art in multiple evaluation metrics.

[0104] Table 1 Quantitative evaluation of different unpaired and zero-shot learning methods on benchmark datasets

[0105]

[0106] "Unpaired set" refers to the set of three unpaired datasets containing DICM, LIME, and MEF.

[0107] On the LOLV1 dataset, the method of the present invention ranks among the top two in the full-reference distortion metrics PSNR and SSIM, demonstrating its effectiveness in structure and detail recovery. Additionally, on the LOLV2 dataset, the method of the present invention outperforms existing zero-shot methods, obtaining the highest PSNR and SSIM scores. For the perceptual metrics on paired datasets including LPIP and FID, the method of the present invention consistently outperforms all methods, ranking first in both evaluations, which highlights its ability to enhance visual quality while retaining perceptual realism. In the unpaired dataset evaluation, the method of the present invention achieves the lowest NIQE and PI scores, indicating excellent perceptual quality in terms of sharpness, contrast, and naturalness. These results show that the method of the present invention effectively balances quantitative accuracy and perceptual quality, demonstrating its superiority and strong generalization ability in real-world scenarios.

[0108] Qualitative comparison

[0109] To facilitate a more intuitive comparison, the present invention presents the visual results of all methods on the paired datasets LOLV1 and LOLV2 in Figure 2 . As observed, the method of the present invention has achieved significant improvements in both color fidelity and brightness balance, thereby achieving a more pleasing enhancement. In contrast, due to over-smoothing or lack of clear constraints and guidance, previous state-of-the-art zero-shot image enhancement methods often have difficulty effectively handling degradation factors, resulting in artifacts and unnatural tones. For example, Clip tends to exhibit insufficient enhancement while producing an over-saturated color effect. Notably, GDP, which utilizes a diffusion prior, faces difficulties in effectively enhancing low-light images in various scenarios. Additionally, Figure 3Illustrates the visual comparison of unpaired datasets. By extracting the illumination prior from the low-frequency wavelet domain, the method of the present invention can preserve more content details while reducing redundant information and noise interference, resulting in a more natural perceptual quality. In contrast, LightEdiverusion introduces a certain degree of color distortion, while GDP and FourierDiff both produce obvious artifacts, further highlighting the robustness and effectiveness of the method of the present invention.

[0110] Ablation study

[0111] Effectiveness of the initial illumination enhancement network.

[0112] As shown in Table 2, the present invention verifies the effectiveness of ILE-NET. Obviously, combining iLenet will have a positive impact on the overall performance. This improvement may stem from the fact that enhancing low-light images in the early stage helps more obvious content extraction and helps obtain the illumination prior for subsequent processing.

[0113] Table 2: Results of the ablation study of the initial illumination enhancement network

[0114]

[0115] The present invention first evaluated the impact of different wavelet transforms on network performance. As shown in Table 3, when no wavelet transform (k = 0) is applied, the network cannot capture the illumination prior from the image, resulting in ineffective enhancement and the worst evaluation performance.

[0116] Table 3 Effectiveness of the number of wavelet transforms K K

[0117] KScale PSNR SSIM LPIPS FID K=0 12.381 0.471 0.439 130.011 K=1 14.809 0.566 0.395 121.501 K=2 19.512 0.751 0.341 81.167 K=3 20.195 0.796 0.301 69.589 K=4 19.701 0.721 0.327 73.167

[0118] As the number of wavelet transforms increases, the present invention observes a significant improvement in overall performance up to k = 3. This confirms that the constructed network can effectively capture the significant light changes in the low-frequency domain after wavelet transformation, and laterally verifies that wavelet transformation can effectively enhance the light in the low-frequency domain. However, when k = 4, the network performance begins to decline. This may be attributed to the substantial change in image size caused by multiple wavelet transforms, which increases the risk of interference from image blur. In addition, excessive light changes at this stage may lead to overexposure. Based on these findings, the present invention sets the default value to k = 3.

[0119] Conclusion

[0120] The present invention proposes a novel zero-shot low-light image enhancement method, which fully utilizes the inherent illumination prior of the image to achieve adaptive enhancement. Through in-depth research on the influence of wavelet transform on illumination changes, the present invention observes that the low-frequency components exhibit significant illumination enhancement characteristics after multiple transformations. Based on this insight, the present invention constructs an initial illumination enhancement network based on the Itinex theory to extract and enhance illumination information. In addition, the present invention designs a transformer-based illumination transfer network, which extracts the illumination background from the low-frequency domain after multiple wavelet transforms and integrates them with the structural content of the initially enhanced image, thereby enhancing adaptively more effectively. Extensive experimental results show that the method of the present invention exhibits excellent interpretability, robustness and efficiency in various low-light scenarios, thus significantly improving the visual quality of low-light images.

[0121] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims concerned.

Claims

1. Zero-shot illumination transfer low-light image enhancement method based on wavelet transform-guided Transformer, characterized in that: At least the following steps are included: S1: Obtain the low-frequency domain of the input low-light image by performing the discrete wavelet transform K times, and jointly feed it into the initial illumination enhancement network to separately extract the reflection map and the illumination map, and restore the initial enhanced image. S2: Subsequently, perform the discrete wavelet transform K times again, extract its low-frequency domain and resize it to obtain a high-brightness illumination image After patch partitioning, both are jointly fed into the Transform Block for feature encoding, and the content structure f C and the illumination background f L are extracted respectively, and then effective feature recombination and fusion are carried out with the help of the Transform Light-Transfer Block; S3: Finally, the final output image is restored through the decoder; S4: In order to further constrain the illumination intensity and effect of the enhanced results, a loss function is used to measure the perceptual effect of the illumination transfer process.

2. The zero-shot wavelet transform-guided Transformer-based illumination transfer low-light image enhancement method according to claim 1, wherein: The initial illumination enhancement network is constructed based on the Retinex theory. The initial illumination enhancement network assumes that the input image I can be decomposed into a reflectance map R and an illumination map L, where R represents the inherent color and texture information of the object, and L reflects the illumination conditions of the image, that is: I=R·L Therefore, the reflection map R and the illumination map L are extracted respectively through two branches, wherein the branch for extracting L uses the input image resized in the low-frequency domain after K discrete wavelet transforms as input, and the subsequent operations are roughly the same in the two branches.

3. The zero-shot wavelet transform-guided Transformer-based illumination transfer low-light image enhancement method according to claim 2, characterized in that: Obtain an enhanced image of the initial restoration Comprising at least the following steps: First, two groups of several convolution blocks are used to obtain embedded features and then processed using the Relu activation function to obtain the reflection feature R1 and the lighting feature L1; Then, the self-attention module is used to refine and further extract the content and illumination information in each branch to obtain the reflection feature R2 and the illumination feature L2; Among them, in the reflection branch, the illumination feature is used to supplement the content information in the reflection feature, while for the illumination branch, the residual structure is used to process the redundant information to obtain clean illumination features; Finally, the processing results of each branch are processed again by convolution blocks, batch normalization and Sigmoid activation function to obtain the final reflectivity map R and illumination map L, which are expressed as: R = Sig(BN(Convs(R2+L2))) L = Sig(BN(Convs(L1-L2))) Among them, Sig represents the Sigmoid activation function; BN represents batch normalization; Convs represents a convolution block composed of multiple convolution operations; Subsequently, based on the obtained R and L, element-wise multiplication is performed in combination with the Retinex theory to obtain a preliminarily restored enhanced image 4. The zero-shot wavelet transform-guided Transformer illumination transfer low-light image enhancement method according to claim 3, wherein: The S2 at least comprises the following steps: First, Perform the discrete wavelet transform K times again, extract its low-frequency domain and resize it to obtain a high-brightness illumination image Subsequently, is used as the content image, and after patch partitioning as the illumination image, they are jointly fed into the TransformBlock for feature encoding to extract content and illumination features respectively; Each Transform Block is a basic box composed of basic LayerNorm, Window Attention and Multilayer Perceptron operations. Specifically: First, the image is partitioned into various non-overlapping n×n patches as tokens in the Transformer, and n is set to 2; The partitioned patch blocks are then sent to the Transform Block for LayerNorm operation, and the window attention module is applied for calculation. Finally, the LayerNorm operation is performed again and the projection mapping is processed through the MLP layer. After each operation, we integrate the residual connection to retain more feature information. The process is formulated as: Among them, Q, K, and V are the query, key, and value matrices respectively; d is the matrix dimension; WA M×M represents multi-head self-attention using a window of shape M×M; B is a bias term; WA 2n×2n represents multi-head self-attention using a window of shape 2n×2n; LN is a layer normalization operation; I l-1 represents the output of the previous layer; MLP represents a multi-layer perceptron; After passing through three layers of Transform Blocks of the same module, content features f are respectively extracted from and ; and lighting features f are extracted from C and L ; The extracted features are then fed into the Transform Light-Transfer Block for integration of content and lighting to learn specific lighting levels for effective enhancement.

5. The zero-shot wavelet transform-guided Transformer-based illumination transfer low-light image enhancement method according to claim 4, characterized in that: The output of S3 is based on the output features with good light enhancement obtained by S1 - S2 in the previous operation. Following the settings of S1 - S2, a Decoder constructed by VGG is used to decode the enhanced features to obtain the final output I out 。 6. The zero-shot wavelet transform-guided Transformer-based illumination transfer low-light image enhancement method according to claim 5, characterized in that: The S4 at least comprises the following steps: Among them, represents the feature map obtained by transforming the input image through a certain layer of feature extraction module; u(·) and respectively represent the mean and variance of the features; At the same time, following previous work, a consistency loss is also adopted to further preserve the structure of the content image and the illumination characteristics of the illumination image, which is formulated as: Among them, I CC represents the output image obtained by extracting light and content from and integrating them; I LL represents the output image obtained by extracting light and content from and integrating them.