Zero-shot low-light image enhancement method based on joint frequency domain prior-guided diffusion

Through the combined frequency domain prior guided diffusion method, low-light images are subjected to frequency domain decomposition and Markov chain structure model processing, combined with wavelet and Fourier transform, the problem of image quality in low-light image enhancement is solved, and the visual effect is improved and the image details are accurately restored.

CN119919293BActive Publication Date: 2025-08-19江苏优众微纳半导体科技有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510108487.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-08-19
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The prior art is difficult to accurately restore image quality in complex and changeable real-life scenarios in low-light image enhancement, and the unsupervised learning method has poor enhancement effect under zero sample conditions, which cannot meet the requirements of image quality consistency and stability.

Method used

Using a method based on joint frequency domain prior guided diffusion, a prior information is constructed for image enhancement by frequency domain decomposition of low-light images, Markov chain structure model and multimodal text supervision, combined with wavelet and Fourier transform.

Benefits of technology

It improves the visual quality and stability of low-light images, accurately restores image details, enhances natural and realistic effects, and is suitable for image reconstruction of complex low-light scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919293B_ABST
    Figure CN119919293B_ABST
Patent Text Reader

Abstract

The present invention discloses a zero-sample low-light image enhancement method based on joint frequency domain prior-guided diffusion, which relates to the field of optics. The method comprises: decomposing the frequency domain of the low-light image; performing forward diffusion and reverse sampling processing on the low-light image using a Markov chain structure model, and combining the results of multimodal text-supervised optimization processing to obtain a sampling result; decomposing the frequency domain of the sampling result to obtain a sampled low-frequency domain; converting the sampled low-frequency domain and low-light low-frequency information using Fourier transform to obtain amplitude and phase information of the sampled low-frequency domain and low-light low-frequency information respectively; combining the amplitude and phase information of the sampled low-frequency domain and low-light low-frequency information with the low-light high-frequency information, and updating the sampling result based on inverse fast Fourier transform and inverse discrete wavelet transform to obtain an image enhancement output. The present invention solves the problem of poor image quality in low-light images by using a joint frequency domain prior-guided diffusion method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of optics, and in particular to a zero-sample low-light image enhancement method based on joint frequency domain prior-guided diffusion. Background Art

[0002] Traditional low-light image enhancement methods primarily rely on optimizing image parameters. While these methods improve image quality to some extent, the hand-crafted prior information lacks adaptability. In complex and variable low-light scenes, it's difficult to accurately and effectively handle varying degrees of insufficient illumination and image degradation, resulting in unstable enhancement results and significant performance variations between images. This makes it difficult to meet the stringent requirements for consistent and stable image quality in practical applications. With the rapid development of deep learning technology, research on low-light image enhancement based on this technology has yielded remarkable results. However, most research focuses excessively on using massive amounts of paired data to fit realistic lighting conditions. In the real world, obtaining large-scale, high-quality paired data is challenging, and model tuning relies heavily on specific datasets, severely limiting the model's generalization capabilities. This means that when faced with the complex and diverse low-light images found in real-world scenarios, the model often struggles to accurately restore the original clarity and rich details, significantly hindering its widespread application in real-world scenarios.

[0003] Given the aforementioned challenges, unsupervised enhancement methods have gradually become a research hotspot. Strategies based on generative models to improve the perceived quality of low-light images have gained some recognition, with the diffusion model standing out. Its superior generative performance has garnered considerable attention and made its mark in the field of supervised image enhancement. However, when applied to unsupervised low-light image enhancement, the lack of prior information about illumination and content significantly reduces its effectiveness in complex and changing real-world scenarios. For example, when processing unknown, severely degraded low-light images, it is difficult to accurately grasp the image's structure and illumination characteristics, resulting in a significant visual gap between the generated image and the natural image, making it impossible to achieve the desired enhancement effect.

[0004] The wavelet domain and the Fourier frequency domain are closely related, bringing new opportunities to solve the problem of zero-sample low-light image enhancement. The wavelet low-frequency domain and Fourier amplitude both gather image illumination information, while the wavelet high-frequency domain and Fourier phase carry image structural information. Moreover, the low-frequency domain after wavelet decomposition can achieve better exposure effects than degraded images, and combined with the Fourier amplitude of normal images, it can accurately guide the restoration of illumination information. In addition, given that high-frequency information is easily damaged during the diffusion process, retaining high-frequency features helps stabilize the image structure and optimize data distribution. On this basis, integrating the wavelet and Fourier frequency domains to construct a diffusion model with rich prior information embedding has become a key breakthrough direction to make up for the lack of illumination and structural information in zero-sample enhancement and improve the stability and quality of the enhancement effect.

[0005] Currently, no effective solutions have been proposed for the problems in related technologies. Summary of the Invention

[0006] In order to overcome the above problems, the present invention aims to propose a zero-sample low-light image enhancement method based on joint frequency domain prior-guided diffusion, with the aim of compensating for the lack of illumination and structural information in zero-sample enhancement and improving the stability and quality of low-light image enhancement effects.

[0007] To this end, the specific technical solutions adopted in the present invention are as follows:

[0008] A zero-sample low-light image enhancement method based on joint frequency domain prior-guided diffusion, the zero-sample low-light image enhancement method based on joint frequency domain prior-guided diffusion includes:

[0009] S1. Decompose the frequency domain of the low-light image to obtain a low-light low-frequency domain, a low-light high-frequency domain, low-light low-frequency information, and low-light high-frequency information;

[0010] S2. Using the Markov chain structure model, forward diffusion and reverse sampling are performed on the low-light image, and the sampling results are obtained by combining the multimodal text supervision optimization processing results;

[0011] S3. Decomposing the frequency domain of the sampling result to obtain a sampled low-frequency domain, converting the sampled low-frequency domain and the low-light low-frequency information using Fourier transform to obtain amplitude and phase information of the sampled low-frequency domain and the low-light low-frequency information respectively;

[0012] S4, combining the amplitude and phase information of the sampled low-frequency domain and the low-light low-frequency information and the low-light high-frequency information, and obtaining an image enhancement output by updating the sampling results based on inverse fast Fourier transform and inverse discrete wavelet transform;

[0013] The S2 includes:

[0014] Input the low-light image into the Markov chain structure model;

[0015] During the forward diffusion process of the low-light image, Gaussian noise is gradually added to the low-light image to obtain a pure noise image;

[0016] During the reverse sampling process of the low-light image, the pure noise image is gradually denoised to obtain a preliminary sampling image.

[0017] Optionally, decomposing the frequency domain of the low-light image to obtain the low-light low-frequency domain, the low-light high-frequency domain, the low-light low-frequency information, and the low-light high-frequency information includes:

[0018] S11. Using discrete wavelet transform, perform diffusion splitting on the low-light image to obtain a low-light low-frequency domain and a low-light high-frequency domain of the low-light image;

[0019] S12. Diffusion splitting is performed on the low-light and low-frequency domain according to discrete wavelet transform to obtain low-light and low-frequency information and low-light and high-frequency information of the low-light and low-frequency domain respectively.

[0020] Optionally, S2 further includes:

[0021] Combining the multimodal model and non-reference brightness control constraints, multimodal text supervision is performed on the preliminary sampling results, and the sampling results are finally optimized.

[0022] Optionally, the forward diffusion process is expressed as:

[0023] ;

[0024] Where, q ( x t | x 0) means that given the initial low-light image x 0 at time t image x t The probability distribution of N represents a normal distribution; x t Indicates at time t Image data; Represents the parameters in the diffusion process, controlling the Gaussian noise added to the image at each moment; Represents the covariance matrix of the Gaussian distribution; I is the identity matrix.

[0025] Optionally, the reverse sampling process is expressed as:

[0026] ;

[0027] Where, Indicates time t Image data Get the moment t -1 Image data The probability distribution of N represents a normal distribution; Indicates the time during reverse sampling t -1 Recovered image data; represents the mean of the Gaussian distribution; Represents the covariance matrix of the Gaussian distribution; I represents the identity matrix; Controls the covariance size.

[0028] Optionally, the preliminary sampling results are supervised by multimodal text by combining the multimodal model and the non-reference brightness control constraint. The final optimized sampling results include:

[0029] Based on the pre-trained multimodal model, the text encoder is fed with preset positive and negative prompts to extract text feature vectors.

[0030] The image encoder is used to process the preliminary sampling results to extract image features, and the similarity loss between the image features and the text feature vector is calculated. The difference between the image features and the text feature vector is determined based on the similarity loss.

[0031] Based on the difference between image features and text feature vectors, the brightness level learning parameters are optimized in combination with non-reference brightness control constraints. The preliminary sampling results are optimized and supervised according to the semantically guided calibration image feature space, and the sampling results are finally optimized.

[0032] Optionally, the expression of the non-reference brightness control constraint is:

[0033] ;

[0034] Where, L bri represents the non-reference brightness control constraint; Indicates the m The average intensity value of non-overlapping local areas; M Indicates the number of different regions in the image; E Indicates the brightness level.

[0035] Optionally, combining the amplitude and phase information of the sampled low-frequency domain and the low-light low-frequency information and the low-light high-frequency information, based on inverse fast Fourier transform and inverse discrete wavelet transform, by updating the sampling results, the image enhancement output includes:

[0036] Combining the wavelet low-frequency domain and Fourier amplitude information to construct brightness prior;

[0037] Combining the high-frequency domain of wavelet and Fourier phase information, the high-frequency domain and phase of the sampling result are replaced by the high-frequency information and phase of the low-light image to construct a priori iterative guided sampling;

[0038] By performing multimodal text supervision on the prior iterative guided sampling, the updated sampling results are obtained;

[0039] Combining the updated sampling results with the low-light and high-frequency domains, an inverse discrete wavelet transform is performed to obtain the image enhancement output.

[0040] Optionally, the a priori iterative guided sampling process is expressed as:

[0041] ;

[0042] Where, represents the prior iterative guided sampling result;IDWT represents the inverse discrete wavelet transform; IFFT represents the inverse fast Fourier transform; represents the brightness learning factor; amp t Indicates the amplitude of the sampled low-frequency domain; amp L Indicates the amplitude of low-light and low-frequency information; pha L Represents the phase of low-light and low-frequency information; Represents low-light high-frequency information.

[0043] Optionally, the expression for the image enhancement output is:

[0044] ;

[0045] Where, I E represents the enhanced output of the image; denoise represents the denoising operation; IDWT represents the inverse discrete wavelet transform; Indicates the updated sampling results; H L Represents the low-light high-frequency domain.

[0046] Compared with the existing technology, the present application has the following beneficial effects: the present invention enhances low-light images by combining Markov chain structure model, discrete wavelet transform, Fourier transform and network optimization. The enhanced low-light images have natural and realistic colors, accurately restore and enrich image details, and the visual presentation is consistent with human eye perception. It can effectively reconstruct the image visual effects in complex low-light scenes, improve visual quality and recognizability, and strongly support its practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The above characteristics, features and advantages of the present invention and their implementation methods and methods will become more clearly understood in conjunction with the following description of the embodiments, which will be described in detail in conjunction with the accompanying drawings. Here, a schematic diagram is shown:

[0048] Figure 1 is a flowchart of a zero-sample low-light image enhancement method based on joint frequency domain prior-guided diffusion according to an embodiment of the present invention;

[0049] Figure 2 is an overall implementation flow chart of a zero-sample low-light image enhancement method based on joint frequency domain prior-guided diffusion according to an embodiment of the present invention;

[0050] Figure 3 This is a visual qualitative comparison diagram of the results of using different enhancement methods in the zero-sample low-light image enhancement method based on joint frequency domain prior-guided diffusion according to an embodiment of the present invention. DETAILED DESCRIPTION

[0051] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0052] According to an embodiment of the present invention, a zero-sample low-light image enhancement method based on joint frequency domain prior-guided diffusion is provided.

[0053] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. Figure 1-Figure 3 As shown, according to an embodiment of the present invention, a zero-sample low-light image enhancement method based on joint frequency domain prior-guided diffusion includes the following steps:

[0054] S1. Decompose the frequency domain of the low-light image to obtain a low-light low-frequency domain, a low-light high-frequency domain, low-light low-frequency information, and low-light high-frequency information.

[0055] Preferably, decomposing the frequency domain of the low-light image to obtain the low-light low-frequency domain, the low-light high-frequency domain, the low-light low-frequency information, and the low-light high-frequency information includes:

[0056] S11. Using discrete wavelet transform, perform diffusion splitting on the low-light image to obtain a low-light low-frequency domain and a low-light high-frequency domain of the low-light image;

[0057] S12. Diffusion splitting is performed on the low-light and low-frequency domain according to discrete wavelet transform to obtain low-light and low-frequency information and low-light and high-frequency information of the low-light and low-frequency domain respectively.

[0058] It should be explained that the input low-light image I L Perform discrete wavelet transform (DWT) diffusion process to low frequency domain, input low light image I L Split into low-light and low-frequency domains L L and low-light high-frequency domain H L , for low light and low frequency domain L L Perform discrete wavelet transform (DWT) again to retain the low-light and low-frequency domains L L Low-light and low-frequency information and low-light high-frequency information , this information will be used to construct the prior and guide the update of the sampling process.

[0059] S2. Use the Markov chain structure model to perform forward diffusion and reverse sampling processing on the low-light image, and combine the multimodal text supervision optimization processing results to obtain the sampling results.

[0060] Preferably, S2 includes:

[0061] Input the low-light image into the Markov chain structure model;

[0062] During the forward diffusion process of the low-light image, Gaussian noise is gradually added to the low-light image to obtain a pure noise image;

[0063] During the reverse sampling process of the low-light image, the pure noise image is gradually denoised to obtain a preliminary sampling image;

[0064] Combining the multimodal model and non-reference brightness control constraints, multimodal text supervision is performed on the preliminary sampling results, and the sampling results are finally optimized.

[0065] Preferably, the forward diffusion process is expressed as:

[0066] ;

[0067] Where, q ( x t | x 0) means that given the initial low-light image x 0 at time t image x t The probability distribution of N represents a normal distribution; x t Indicates at time t Image data; Represents the parameters in the diffusion process, controlling the Gaussian noise added to the image at each moment; Represents the covariance matrix of the Gaussian distribution; I is the identity matrix.

[0068] Preferably, the reverse sampling process is expressed as:

[0069] ;

[0070] Where, Indicates time t Image data Get the moment t -1 Image data The probability distribution of N represents a normal distribution; Indicates the time during reverse sampling t -1 Recovered image data; represents the mean of the Gaussian distribution; Represents the covariance matrix of the Gaussian distribution; I represents the identity matrix; Controls the covariance size.

[0071] Preferably, the multimodal model and the non-reference brightness control constraint are combined to perform multimodal text supervision on the preliminary sampling results, and the final optimized sampling results include:

[0072] Based on the pre-trained multimodal model, the text encoder is fed with preset positive and negative prompts to extract text feature vectors.

[0073] The image encoder is used to process the preliminary sampling results to extract image features, and the similarity loss between the image features and the text feature vector is calculated. The difference between the image features and the text feature vector is determined based on the similarity loss.

[0074] Based on the difference between image features and text feature vectors, the brightness level learning parameters are optimized in combination with non-reference brightness control constraints. The preliminary sampling results are optimized and supervised according to the semantically guided calibration image feature space, and the sampling results are finally optimized.

[0075] Preferably, the expression of the non-reference brightness control constraint is:

[0076] ;

[0077] Where, L bri represents the non-reference brightness control constraint; Indicates the m The average intensity value of non-overlapping local areas; M Indicates the number of different regions in the image; E Indicates the brightness level.

[0078] It needs to be explained that the forward propagation process of the Markov chain structure q By adding x 0 is achieved by gradually adding Gaussian noise until T time steps approximate pure noise data x T . Forward propagation process q The noise in is independent and follows a normal distribution. The expression of the forward diffusion process is:

[0079] ;

[0080] Where, q (x t | x 0) means that given the initial low-light image x 0 at time t image x t The probability distribution of N represents a normal distribution; x t Indicates at time t Image data; Represents the parameters in the diffusion process, controlling the Gaussian noise added to the image at each moment. , a t =1- β t , t ∈[1,…, T ], β t Used to determine the parameters for adding Gaussian noise during the forward diffusion process; Represents the covariance matrix of the Gaussian distribution; I is the identity matrix.

[0081] Reverse sampling process p θ Mainly for pure noise x T Denoising to restore the image x 0, the expression of the reverse sampling process is:

[0082] ;

[0083] Where, Indicates time t Image data Get the moment t -1 Image data The probability distribution of N represents a normal distribution; Indicates the time during reverse sampling t -1 Recovered image data; represents the mean of the Gaussian distribution, where represents the estimated noise; Represents the covariance matrix of the Gaussian distribution; I represents the identity matrix; Indicates controlling the covariance size.

[0084] Combining multimodal models and non-reference brightness control constraints, multimodal text supervision is performed on the preliminary sampled images, and a pre-trained CLIP model is introduced to input preset positive prompts into the text encoder. Tp and negative prompts T n Extract feature vectors and use image encoder to sample the results of each step x t Extract features and then calculate the similarity loss between image and text vectors in CLIP latent space L TG ,According to this supervised enhancement process, the image feature space is aligned with ,the semantic guidance capability to improve the image perception quality.

[0085] The expression of similarity loss LTG is:

[0086] ;

[0087] Where LTG represents the similarity loss between image and text vector; Φ image Represents the image encoder, which is used to sample the results x t Perform feature extraction and convert image data into feature vectors; Φ text Represents a text encoder, which is used to extract features from preset text prompts and convert text information into feature vectors; T p Indicates that the preset is prompting; T n Indicates preset negative prompt; T j Represents a text prompt; j The number representing the sample; e cos An index representing the cosine similarity between image features and text features.

[0088] Introducing non-reference brightness control constraints L bri , which can optimize the brightness level learning parameters , which can ensure good visual perception and brightness distribution of the sampling results and enhance the stability and reliability of the image enhancement effect. Non-reference brightness control constraints L bri The expression is:

[0089] ;

[0090] Where, L bri represents the non-reference brightness control constraint; Indicates the m The average intensity value of non-overlapping local areas; M Indicates the number of different regions in the image; E Indicates the brightness level.

[0091] S3. Decompose the frequency domain of the sampling result using discrete wavelet transform to obtain the sampling low-frequency domain of the sampling result. According to the sampling low-frequency domain and the low-light low-frequency information, use Fourier transform to obtain the amplitude and phase information of the sampling low-frequency domain and the low-light low-frequency information respectively.

[0092] It should be explained that the expressions for the amplitude and phase information of the sampled low-frequency domain and low-light low-frequency information are:

[0093] ;

[0094] ;

[0095] Where, amp t Indicates the amplitude of the sampled low-frequency domain; pha t Represents the phase of the low-light low-frequency domain; amp L Indicates the amplitude of low-light and low-frequency information; pha L Represents the phase of low-light and low-frequency information; FFT represents the Fourier transform.

[0096] S4. Combining the amplitude and phase information of the sampled low-frequency domain and the low-light low-frequency information with the low-light high-frequency information, based on inverse fast Fourier transform and inverse discrete wavelet transform, the image enhancement output is obtained by updating the sampling results.

[0097] Preferably, combining the amplitude and phase information of the sampled low-frequency domain and the low-light low-frequency information and the low-light high-frequency information, based on inverse fast Fourier transform and inverse discrete wavelet transform, by updating the sampling results, the image enhancement output includes:

[0098] Combining the wavelet low-frequency domain and Fourier amplitude information to construct brightness prior;

[0099] Combining the high-frequency domain of wavelet and Fourier phase information, the high-frequency domain and phase of the sampling result are replaced by the high-frequency information and phase of the low-light image to construct a priori iterative guided sampling;

[0100] By performing multimodal text supervision on the prior iterative guided sampling, the updated sampling results are obtained;

[0101] Combining the updated sampling results with the low-light and high-frequency domains, an inverse discrete wavelet transform is performed to obtain the image enhancement output.

[0102] Preferably, the expression of the a priori iterative guided sampling process is:

[0103] ;

[0104] Where, represents the prior iterative guided sampling result; IDWT represents the inverse discrete wavelet transform; IFFT represents the inverse fast Fourier transform; represents the brightness learning factor; amp t Indicates the amplitude of the sampled low-frequency domain; amp L Indicates the amplitude of low-light and low-frequency information; pha L Represents the phase of low-light and low-frequency information; Represents low-light high-frequency information.

[0105] Preferably, the expression of the image enhancement output is:

[0106] ;

[0107] Where, I E represents the enhanced output of the image; denoise represents the denoising operation; IDWT represents the inverse discrete wavelet transform; Indicates the updated sampling results; H L Represents low-light high-frequency domain

[0108] It should be explained that, since the low-frequency domain of wavelet and Fourier amplitude are closely related to the brightness and structural information of the image respectively, the two are combined to construct the brightness prior. Considering that the high-frequency information of wavelet and the sampling phase are prone to produce random details that interfere with the distribution of image data, the low-light high-frequency information of low-light images is used. and phase pha L Instead, it guides the generation of sampling content and maintains the stability of data distribution, ensuring the fidelity and consistency of data distribution.

[0109] like Figure 2 The overall implementation flow chart of this method is shown in Figure 2. During the experimental simulation, the experimental framework was built on a single NVIDIA Tesla V100 GPU based on the PyTorch framework. The unconditional 256×256 diffusion model pre-trained on ImageNet was used. The pre-trained weights learned on large-scale image data gave the model excellent initial feature extraction and expression capabilities, accelerating the convergence of low-light image enhancement tasks and improving generalization. The total diffusion step size was set to TThe value of 1000 is precisely set to strictly control the gradual transition of the image from its original pure state to completely submerged in noise. The alternating optimization interval S is set to 200, and optimization adjustments are performed every 200 steps during inverse sampling, balancing computational cost and model learning efficiency. This ensures precise optimization of model parameters in the key stages of denoising, lighting compensation, and detail restoration, allowing the generated image to gradually approximate the distribution of natural image data.

[0110] In order to verify the effectiveness of the present invention, test images were selected from the LOL (Deep Retinex Decomposition for Low-Light Enhancement) and SICE (Single Image Contrast Enhancement) paired datasets, and low-light images were collected from the LIME (Low-light Image Enhancement via Illumination Map Estimation), DICM (Contrast enhancement based on layered difference representation), and MEF (Power-constrained contrast enhancement for emissive displays based on histogram equalization) datasets to form an "Unpaired - set" (representing an unpaired dataset). The paired dataset is evaluated using PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index), LPIPS (Learned Perceptual Image Patch Similarity), and FID (Frechet Inception Distance). The unpaired dataset is evaluated using MUSIQ (Multi-Scale Image Quality Assessment) and LOE (Low-Light Enhancement).

[0111] Table 1 below shows a quantitative evaluation of different unsupervised learning methods on benchmark datasets. Among these, the zero-shot low-light image enhancement method based on joint frequency-domain prior-guided diffusion achieves outstanding overall performance among zero-shot methods, achieving a PSNR of 20.922, an SSIM of 0.811, a low LPIPS of 0.281, and a FID of 63.601. This method demonstrates the effectiveness and strong generalization of the model, and precise quantitative analysis demonstrates its significant advantages in improving image quality, optimizing structural similarity, reducing perceptual differences, and aligning with data distribution.

[0112] Table 1 Quantitative evaluation of different unsupervised learning methods on benchmark datasets

[0113]

[0114] like Figure 3 Figure 2 shows a qualitative visual comparison of the results of different enhancement methods. Compared to other models, the proposed method achieves natural and realistic colors, accurately restores rich image details, and presents visual presentation consistent with human perception. It can effectively reconstruct image visual effects in complex low-light scenes, improving visual quality and legibility, strongly supporting its practical application value.

[0115] Specifically, in order to facilitate better understanding by those skilled in the art, the relevant embodiments of the present application will now explain the technical terms or some nouns that may be involved in the present application.

[0116] To sum up, with the help of the above-mentioned technical solution of the present invention, by combining the Markov chain structure model, discrete wavelet transform, Fourier transform and network optimization, low-light images are enhanced. The enhanced low-light images have natural and realistic colors, accurately restore and enrich the image details, and the visual presentation is consistent with the human eye perception. It can effectively reconstruct the image visual effects in complex low-light scenes, improve visual quality and recognizability, and strongly support its practical application value.

[0117] Although the present invention has been disclosed above with reference to preferred embodiments, the embodiments are merely examples for the purpose of illustration and are not intended to limit the present invention. Those skilled in the art may make various modifications and alterations without departing from the spirit and scope of the present invention. The scope of protection claimed by the present invention shall be subject to the claims.

Claims

1. A zero-shot low-light image enhancement method based on joint frequency domain prior-guided diffusion, characterized by: include: S1. Decompose the frequency domain of the low-light image to obtain a low-light low-frequency domain, a low-light high-frequency domain, low-light low-frequency information, and low-light high-frequency information; S2. Using the Markov chain structure model, forward diffusion and reverse sampling are performed on the low-light image, and the sampling results are obtained by combining the multimodal text supervision optimization processing results; S3. Decomposing the frequency domain of the sampling result to obtain a sampled low-frequency domain, converting the sampled low-frequency domain and the low-light low-frequency information using Fourier transform to obtain amplitude and phase information of the sampled low-frequency domain and the low-light low-frequency information respectively; S4, combining the amplitude and phase information of the sampled low-frequency domain and the low-light low-frequency information and the low-light high-frequency information, and obtaining an image enhancement output by updating the sampling results based on inverse fast Fourier transform and inverse discrete wavelet transform; The S2 includes: Input the low-light image into the Markov chain structure model; During the forward diffusion process of the low-light image, Gaussian noise is gradually added to the low-light image to obtain a pure noise image; The expression of the forward diffusion process is: ; Where, q ( x t | x 0) means that given the initial low-light image x 0 at time t image x t The probability distribution of N represents a normal distribution; x t Indicates at time t Image data; Represents the parameters in the diffusion process, controlling the Gaussian noise added to the image at each moment; Represents the covariance matrix of the Gaussian distribution; I is the identity matrix; During the reverse sampling process of the low-light image, the pure noise image is gradually denoised to obtain a preliminary sampling image; The expression of the reverse sampling process is: ; Where, Indicates time t Image data Get the moment t -1 Image data The probability distribution of N represents a normal distribution; Indicates the time during reverse sampling t -1 Recovered image data; represents the mean of the Gaussian distribution; Represents the covariance matrix of the Gaussian distribution; I represents the identity matrix; Controls the covariance size.

2. The zero-shot low-light image enhancement method based on joint frequency domain prior-guided diffusion according to claim 1, characterized in that: Decomposing the frequency domain of the low-light image to obtain the low-light low-frequency domain, the low-light high-frequency domain, the low-light low-frequency information, and the low-light high-frequency information includes: S11. Using discrete wavelet transform, perform diffusion splitting on the low-light image to obtain a low-light low-frequency domain and a low-light high-frequency domain of the low-light image; S12. Diffusion splitting is performed on the low-light and low-frequency domain according to discrete wavelet transform to obtain low-light and low-frequency information and low-light and high-frequency information of the low-light and low-frequency domain respectively.

3. The zero-shot low-light image enhancement method based on joint frequency domain prior-guided diffusion according to claim 1, characterized in that: Said S2 further comprises: Combining the multimodal model and non-reference brightness control constraints, multimodal text supervision is performed on the preliminary sampling results, and the sampling results are finally optimized.

4. The zero-shot low-light image enhancement method based on joint frequency domain prior-guided diffusion according to claim 3, characterized in that: The multimodal model and the non-reference brightness control constraint are combined to perform multimodal text supervision on the preliminary sampling results, and the final optimized sampling results include: Based on the pre-trained multimodal model, the text encoder is fed with preset positive and negative prompts to extract text feature vectors. The image encoder is used to process the preliminary sampling results to extract image features, and the similarity loss between the image features and the text feature vector is calculated. The difference between the image features and the text feature vector is determined based on the similarity loss. Based on the difference between image features and text feature vectors, the brightness level learning parameters are optimized in combination with non-reference brightness control constraints. The preliminary sampling results are optimized and supervised according to the semantically guided calibration image feature space, and the sampling results are finally optimized.

5. The zero-shot low-light image enhancement method based on joint frequency domain prior-guided diffusion according to claim 4, characterized in that: The expression of the non-reference brightness control constraint is: ; Where, L bri represents the non-reference brightness control constraint; Indicates the m The average intensity value of non-overlapping local areas; M Indicates the number of different regions in the image; E Indicates the brightness level.

6. The zero-shot low-light image enhancement method based on joint frequency domain prior-guided diffusion according to claim 1, characterized in that: The image enhancement output obtained by combining the amplitude and phase information of the sampled low-frequency domain and the low-light low-frequency information and the low-light high-frequency information based on the inverse fast Fourier transform and the inverse discrete wavelet transform by updating the sampling results includes: Combining the wavelet low-frequency domain and Fourier amplitude information to construct brightness prior; Combining the high-frequency domain of wavelet and Fourier phase information, the high-frequency domain and phase of the sampling result are replaced by the high-frequency information and phase of the low-light image to construct a priori iterative guided sampling; By performing multimodal text supervision on the prior iterative guided sampling, the updated sampling results are obtained; Combining the updated sampling results with the low-light and high-frequency domains, an inverse discrete wavelet transform is performed to obtain the image enhancement output.

7. The zero-shot low-light image enhancement method based on joint frequency domain prior-guided diffusion according to claim 6, characterized in that: The expression of the a priori iterative guided sampling process is: ; Where, represents the prior iterative guided sampling result; IDWT represents the inverse discrete wavelet transform; IFFT represents the inverse fast Fourier transform; represents the brightness learning factor; amp t Indicates the amplitude of the sampled low-frequency domain; amp L Indicates the amplitude of low-light and low-frequency information; pha L Represents the phase of low-light and low-frequency information; Represents low-light high-frequency information.

8. The zero-shot low-light image enhancement method based on joint frequency domain prior-guided diffusion according to claim 6, characterized in that: The expression of the image enhancement output is: ; Where, I E represents the enhanced output of the image; denoise represents the denoising operation; IDWT represents the inverse discrete wavelet transform; Indicates the updated sampling results; H L Represents the low-light high-frequency domain.

Citation Information

Patent Citations

  • Low-rank recovery and deep diffusion fusion-based low-light image enhancement method

    CN119338730A