A Low-Light Image Enhancement Method Based on the Fusion of Low-Rank Recovery and Deep Diffusion
By integrating low-rank recovery and depth diffusion technology in the low-light image enhancement algorithm, and using the global depth diffusion generation module and the discrete cosine transformation module for image processing, the problem of noise trace residue during image denoising in the prior art is solved, and a higher quality low-light image enhancement effect is achieved.
Patent Information
- Application Number
- CN202411885652.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-20
AI Technical Summary
The existing low-light image enhancement algorithm based on diffusion model is difficult to effectively capture the texture and detail characteristics of the image during denoising processing, resulting in obvious noise traces in the enhanced image, affecting clarity and quality.
Using a method based on low-rank recovery and deep diffusion fusion fusion, a global deep diffusion generation module is formed by adding a Markov chain mechanism before the Transformer network, combining the discrete cosine transformation module and the inverse transformation module, forward noise addition and denoising processing is carried out, the global characteristics of the noise distribution are captured, the noise components are gradually removed, and the clarity and quality of the image are improved.
By fusing low-rank recovery and depth diffusion techniques, the clarity and quality of low-light images are significantly improved, noise traces are reduced, and the visual effect of the image is enhanced.
Smart Images

Figure CN119338730B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a low-light image enhancement method based on the fusion of low-rank recovery and deep diffusion. Background Art
[0002] For images obtained under low-light conditions, their visual quality is often greatly reduced. This not only makes the images look dim and blurred, but also results in a large amount of loss of detail information and a significant increase in noise. Such image quality undoubtedly poses great challenges to subsequent downstream vision tasks such as image processing, feature extraction, and target recognition. Therefore, enhancing low-light images to improve their visual quality is the key to ensuring the accurate and efficient execution of downstream vision tasks.
[0003] In the prior art, diffusion models have received attention due to their performance in image generation and restoration tasks. First, the diffusion model analyzes the noise distribution characteristics in low-light images. By accurately modeling the noise, it grasps the noise components in low-light images. Subsequently, using the reverse diffusion process, it gradually guides the image to transform from a noisy low-light state to a clear and bright state. During the transformation process, the noise in the image is gradually eliminated, and the details, brightness, and contrast of the image are gradually restored and enhanced using the latent information and structural features in the image to obtain an enhanced image.
[0004] The defects of the above prior art are that when the low-light image enhancement algorithm based on the diffusion model performs noise addition and noise removal processing on low-light images, the capture of texture and detail features mainly focuses on the spatial domain, resulting in obvious noise traces remaining on the denoised image, thereby affecting the clarity and quality of the enhanced image. Summary of the Invention
[0005] Based on this, it is necessary to provide a low-light image enhancement method based on the fusion of low-rank recovery and deep diffusion for the above technical problems.
[0006] An embodiment of the present invention provides a low-light image enhancement method based on the fusion of low-rank recovery and deep diffusion, including:
[0007] Obtain historical low-light images and corresponding historical high-light images;
[0008] Add a Markov chain mechanism before the Transformer network to form a global deep diffusion generation module. The input end of the global deep diffusion generation module is connected to the output end of the discrete cosine transform module, and the output end of the global deep diffusion generation module is connected to the input end of the inverse discrete cosine transform module. Take the input end of the discrete cosine transform module as the input end of the low-light image enhancement model, and take the output end of the inverse discrete cosine transform module as the output end of the low-light image enhancement model to obtain a low-light image enhancement model;
[0009] Input the historical low-light image into the low-light image enhancement model to obtain the enhanced result of the historical low-light image; train the low-light image enhancement model based on the enhanced result of the historical low-light image and the corresponding historical high-light image of the historical low-light image.
[0010] Input the low-light image to be processed into the trained low-light image enhancement model, transform the pixel values of the low-light image to be processed through the discrete cosine transform module to obtain a spectrogram in the frequency domain; perform forward noise addition on the spectrogram at each time step through the Markov chain mechanism in the global depth diffusion generation module, and then capture the global features of the noise distribution through the Transformer network in the global depth diffusion generation module; in each time step, remove the noise components in the current time step according to the global features of the noise distribution to obtain an intermediate image; use the intermediate image corresponding to the last time step as the new spectrogram, and perform inverse discrete cosine transform on the new spectrogram through the inverse discrete cosine transform module to convert the new spectrogram from the frequency domain back to the spatial domain to obtain the enhanced low-light image.
[0011] Optionally, transforming the pixel values of the low-light image to be processed through the discrete cosine transform module specifically includes:
[0012] The discrete cosine transform decomposes the pixel values of the low-light image to be processed, expressed as a weighted sum of different frequency components;
[0013] The formula for the transformation process:
[0014]
[0015] where D(u, v) represents the coefficient in the frequency domain, I(x, y) represents the pixel value of the original image, x and y represent the coordinates of the image in the horizontal and vertical directions respectively, u and v represent frequency variables, N represents the size of the image, and α(u)α(v) represents the normalization factor;
[0016] The definition formula of the normalization factor is:
[0017]
[0018] where i represents a formal parameter, and numerically equals u and v.
[0019] Optionally, perform forward noise addition on the spectrogram at each time step through the Markov chain mechanism in the global depth diffusion generation module, and the formula for forward noise addition is:
[0020]
[0021] where t represents the time step, T represents the last time step, x 0 ,x1 ,..., x T represents the images generated at each step, x 0 represents the original image, x T represents pure Gaussian noise, ò represents Gaussian noise satisfying the (0,1) distribution, represents a variable that becomes smaller as the time step increases;
[0022] The relationship between the variables and hyperparameters is:
[0023] α t = 1 - β t
[0024]
[0025] where, α t represents the variable at time step t, β t represents a hyperparameter that becomes larger as the time step increases.
[0026] Optionally, the global features of the noise distribution are captured by the Transformer network in the global depth diffusion generation module, and the calculation formula for the global features is:
[0027]
[0028] where, Q represents the local feature matrix after dynamic convolution processing, Q T represents the transpose matrix of the local feature matrix Q, d k represents the dimension of the global feature K.
[0029] Optionally, the noise components in the current time step are removed according to the global features of the noise distribution, and the formula for the removal process is:
[0030]
[0031] where, x t represents the sample at time step t in the diffusion process, ò θ (x t , t) represents the noise prediction function of the model, σ t represents the coefficient that controls the amount of noise at time step t, z represents random noise from the standard normal distribution, represents a variable that becomes smaller as the time step increases, α t represents the variable at time step t.
[0032] Optionally, the inverse discrete cosine transform is performed on the new spectrogram through the inverse discrete cosine transform module, specifically including:
[0033] The frequency domain information is converted back to the spatial domain through the inverse discrete cosine transform, and the image is reconstructed into a visible form;
[0034] The inverse discrete cosine transform formula is expressed as:
[0035]
[0036] Among them, D(u, v) represents the coefficient in the frequency domain, I(x, y) represents the pixel value of the original image, x and y respectively represent the coordinates of the image in the horizontal and vertical directions, u and v represent the frequency variables, N represents the size of the image, and α(u)α(v) represents the normalization factor.
[0037] Compared with the prior art, the beneficial effects of the above-mentioned low-light image enhancement method based on the fusion of low-rank recovery and depth diffusion provided by the embodiments of the present invention are as follows:
[0038] In the present invention, the discrete cosine transform module is used to transform the pixel values of the low-light image to be processed, so that the structural information of the low-light image to be processed can be more clearly analyzed in the frequency domain, effectively improving the generation efficiency; the Markov chain mechanism in the global depth diffusion generation module adds noise forward at each time step, and then the Transformer network in the global depth diffusion generation module captures the global features of the noise distribution, and removes the noise components in the current time step according to the global features of the noise distribution to obtain an intermediate image; the intermediate image corresponding to the last time step is used as a new spectrogram, and the global nature of the noise components is emphasized in this process, which can improve the clarity and quality of the enhanced low-light image. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 FIG. is a processing flow chart of a low-light image enhancement method based on the fusion of low-rank recovery and depth diffusion provided in an embodiment;
[0040] Figure 2 FIG. is a schematic diagram of the action of discrete cosine transform of a low-light image enhancement method based on the fusion of low-rank recovery and depth diffusion provided in an embodiment;
[0041] Figure 3 FIG. is a convergence graph of the loss function of a low-light image enhancement method based on the fusion of low-rank recovery and depth diffusion provided in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0043] In one embodiment, a low-light image enhancement method based on the fusion of low-rank recovery and depth diffusion is provided, and the method includes:
[0044] Obtain the historical low-light image and the corresponding historical high-light image of the historical low-light image.
[0045] Add a Markov chain mechanism before the Transformer network to form a global depth diffusion generation module. The input end of the global depth diffusion generation module is connected to the output end of the discrete cosine transform module, and the output end of the global depth diffusion generation module is connected to the input end of the inverse discrete cosine transform module. Take the input end of the discrete cosine transform module as the input end of the low-light image enhancement model, and take the output end of the inverse discrete cosine transform module as the output end of the low-light image enhancement model to obtain the low-light image enhancement model, as Figure 1 shown.
[0046] Input the historical low-light image into the low-light image enhancement model to obtain the enhancement result of the historical low-light image. Train the low-light image enhancement model based on the enhancement result of the historical low-light image and the corresponding historical high-light image of the historical low-light image.
[0047] Input the low-light image to be processed into the trained low-light image enhancement model. Transform the pixel values of the low-light image to be processed through the discrete cosine transform module to obtain a spectrogram in the frequency domain. Forward add noise to the spectrogram at each time step through the Markov chain mechanism in the global depth diffusion generation module, and then capture the global features of the noise distribution through the Transformer network in the global depth diffusion generation module. At each time step, remove the noise components in the current time step according to the global features of the noise distribution to obtain an intermediate image. Take the intermediate image corresponding to the last time step as the new spectrogram, and perform inverse discrete cosine transform on the new spectrogram through the inverse discrete cosine transform module to convert the new spectrogram from the frequency domain back to the spatial domain to obtain the enhanced low-light image.
[0048] Provide a specific embodiment of the present invention:
[0049] S1. Construct a low-light image enhancement model based on the fusion of low-rank recovery and depth diffusion; wherein the low-light image enhancement model includes a discrete cosine transform module, a global depth diffusion generation module, and an inverse discrete cosine transform module;
[0050] S2. Preprocess the low-light images and normal images in the dataset, including resizing, normalizing, etc., to ensure the quality and consistency of the input data. Send the preprocessed pictures into the discrete cosine transform module to obtain the spectrogram corresponding to each picture. Each point in the spectrogram represents the components of different frequencies in the image. As Figure 2As shown, this module converts the information in the spatial domain of the picture into the frequency domain, performs high-frequency truncation to reduce high-frequency data, and finally uses the inverse discrete cosine transform to convert the information back to the spatial domain, realizing the conversion of information between the spatial domain and the frequency domain.
[0051] Specifically, the discrete cosine transform decomposes the pixel values of the low-light image to be processed, expressed as a weighted sum of different frequency components. In this way, the structural information of the image can be more clearly analyzed in the frequency domain. The formula for the discrete cosine transform is:
[0052]
[0053] where D(u, v) represents the coefficient in the frequency domain, I(x, y) represents the pixel value of the original image, x and y represent the coordinates of the image in the horizontal and vertical directions respectively, u and v represent the frequency variables, N represents the size of the image, and α(u)α(v) represents the normalization factor, defined as:
[0054]
[0055] where i is a formal parameter for representation, and numerically equals u and v.
[0056] Specifically, the data matrix of natural images is usually low-rank or approximately low-rank, and most of the energy is concentrated in a few low-frequency components, which reflect the main features of the image. However, the noise and errors generated during the acquisition or processing process appear in the form of high-frequency components, destroying the low-rank property of the image. To restore the low-rank property of the image and denoise, high-frequency truncation can be performed in the frequency domain, retaining the low-frequency components after the DCT transform, thereby effectively removing noise and retaining the core features, and effectively accelerating the training and inference process.
[0057] S3. Input the spectrogram obtained in S2 into the diffusion model, perform forward noise addition using the mechanism of the Markov chain. At each time step, the model gradually adds noise to the spectrogram until it approaches Gaussian noise. Use the Transformer network to learn the noise distribution at each time step, capture the global features of the noise distribution, and extract relevant information from the noise through the self-attention mechanism, so as to better simulate the noise characteristics of real data at each time step.
[0058] Specifically, the spectrogram after the discrete cosine transform is used to perform forward noise addition using the mechanism of the Markov chain. At each time step, the model gradually adds noise to the spectrogram to generate a noisy sample, and its specific formula is as follows:
[0059]
[0060] where t represents the time step, T represents the last time step, x0 , x 1 ,..., x T represents the images generated at each step, x 0 represents the original image, x T represents pure Gaussian noise, and ò represents Gaussian noise satisfying the (0, 1) distribution. represents a variable that becomes smaller as the time step increases.
[0061] The relationship between the variable and the hyperparameter is:
[0062] α t = 1 - β t
[0063]
[0064] Among them, α t represents the variable at time step t, and β t represents a hyperparameter that becomes larger as the time step increases.
[0065] Meanwhile, the Transformer network obtains global information through the residual self-attention mechanism. This mechanism enhances the weight of the original features and the results of the previous layer's attention mechanism in subsequent attention operations, making the global information more robust. Incorporating it into the global depth diffusion generation module, the formula is as follows:
[0066]
[0067] Among them, Q represents the local feature matrix after dynamic convolution processing, and Q T represents the transpose matrix of the local feature matrix Q, and d k represents the dimension of the global feature K.
[0068] S4. Initialize a noise combined with the low-light image. Using the noise distribution learned by the Transformer network, at each time step t, the model refers to the current input noise image and predicts and removes the noise component of the current step according to the learned noise distribution. This denoising process is carried out step by step. The model gradually removes the noise through a series of time steps, generating a clearer intermediate image at each step until a noise-free spectrogram approaching the real data is restored.
[0069] Specifically, it includes: removing the noise component in the current time step according to the global feature of the noise distribution, and its calculation formula is:
[0070]
[0071] Among them, x t represents the sample at time step t in the diffusion process, and ò θ (xt , t) represents the noise prediction function of the model, σ t represents the coefficient that controls the amount of noise at time step t, z represents the random noise from the standard normal distribution, represents a variable that becomes smaller as the time step increases, α t represents the variable at time step t.
[0072] S5. Perform the inverse discrete cosine transform on the spectrogram to convert the frequency domain information back to the spatial domain, enabling the image to be reconstructed into a visible form. This method improves the clarity and visual quality of the image, making the restored high-light image more realistic and natural.
[0073] Specifically, by performing the inverse discrete cosine transform on the spectrogram, the frequency domain information can be effectively converted back to the spatial domain. Since the spectral image only contains frequency components and lacks direct visible information. The implementation of the inverse discrete cosine transform not only allows the reconstruction of the visible form of the image but also ensures the integrity of its main features. The specific formula is as follows:
[0074]
[0075] where D(u, v) represents the coefficient in the frequency domain, I(x, y) represents the pixel value of the original image, x and y represent the coordinates of the image in the horizontal and vertical directions respectively, similar to the case in the discrete spectral transform, N represents the size of the image, and α(u)α(v) represents the normalization factor.
[0076] In this process, the inverse discrete cosine transform recombines all frequency components, enabling each pixel value to reflect the corresponding frequency information. This transformation not only restores the visibility of the image but also maintains the overall quality and details of the image, making the generated image visually clearer, more natural, and rich in details.
[0077] Exemplarily, when providing the algorithm based on the diffusion model and training it on the low-light LOL dataset with the present invention, the change curve of the loss function, and the convergence graph is as Figure 3 shown. It can be seen that the present invention has a faster convergence speed of the loss function and a smaller training loss, indicating that the present invention can effectively improve the model training speed and accuracy.
[0078] The above-described embodiments only represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A low-light image enhancement method based on low-rank recovery and deep diffusion fusion, characterized in that: include: Obtain historical low-light images and historical high-light images corresponding to the historical low-light images; A Markov chain mechanism is added before the Transformer network to form a global depth diffusion generation module. The input end of the global depth diffusion generation module is connected to the output end of the discrete cosine transform module, and the output end of the global depth diffusion generation module is connected to the input end of the inverse discrete cosine transform module. The input end of the discrete cosine transform module is used as the input end of the low-light image enhancement model, and the output end of the inverse discrete cosine transform module is used as the output end of the low-light image enhancement model to obtain the low-light image enhancement model. Input the historical low-light image into the low-light image enhancement model to obtain the enhancement result of the historical low-light image; Based on the enhancement results of historical low-light images and the historical high-light images corresponding to the historical low-light images, a low-light image enhancement model is trained; The low-light image to be processed is input into the trained low-light image enhancement model, and the pixel values of the low-light image to be processed are transformed through the discrete cosine transform module to obtain a spectrum map in the spectral domain; the spectrum map is forward denoised at each time step through the Markov chain mechanism in the global deep diffusion generation module, and then the global characteristics of the noise distribution are captured through the Transformer network in the global deep diffusion generation module; in each time step, the noise component in the current time step is removed according to the global characteristics of the noise distribution to obtain an intermediate image; the intermediate image corresponding to the last time step is used as the new spectrum map, and the new spectrum map is inversely discretely transformed through the inverse discrete cosine transform module to convert the new spectrum map from the spectral domain back to the spatial domain to obtain an enhanced low-light image.
2. The low-light image enhancement method based on low-rank restoration and deep diffusion fusion as claimed in claim 1, characterized in that: The transforming of the pixel values of the low-light image to be processed by the discrete cosine transform module specifically includes: Discrete cosine transform decomposes the pixel values of the low-light image to be processed into a weighted sum of different frequency components; The formula for the transformation process is: ; in, represents the coefficients in the frequency domain, Represents the pixel value of the original image, x and y Respectively represent the horizontal and vertical coordinates of the image, and represents a frequency variable, N Indicates the size of the image. represents the normalization factor; The definition of the normalization factor is: ; in, i Represents a formal parameter, numerically equal to u and v .
3. The low-light image enhancement method based on low-rank restoration and deep diffusion fusion as claimed in claim 1, characterized in that: The Markov chain mechanism in the global deep diffusion generation module performs forward noise addition on the spectrum graph at each time step. The formula for forward noise addition is: ; in, t represents the time step, T represents the last time step, Represents the pictures generated at each step, Represents the original image, represents pure Gaussian noise, represents Gaussian noise that satisfies the (0,1) distribution, represents a variable that gets smaller as the time step increases; The relationship between variables and hyperparameters is: ; in, Represents the time step t Variables when represents a hyperparameter that gets larger as the number of time steps increases.
4. The low-light image enhancement method based on low-rank restoration and deep diffusion fusion as claimed in claim 1, characterized in that: The global features of the noise distribution are captured by the Transformer network in the global deep diffusion generation module. The calculation formula of the global features is: ; in, represents the local feature matrix after dynamic convolution processing, Represents the local feature matrix The transposed matrix of Represents global features k The dimension of .
5. The low-light image enhancement method based on low-rank restoration and deep diffusion fusion as claimed in claim 1, characterized in that: The noise component in the current time step is removed according to the global characteristics of the noise distribution, and the formula of the removal process is: ; in, represents the time step in the diffusion process t The sample at represents the noise prediction function of the model, Represents the time step t The coefficient that controls the amount of noise, represents random noise from a standard normal distribution, represents a variable that gets smaller as the time step increases, Represents the time step t Variables when .
6. The low-light image enhancement method based on low-rank restoration and deep diffusion fusion as claimed in claim 1, characterized in that: The inverse discrete cosine transform of the new spectrum graph is performed by the inverse discrete cosine transform module, specifically comprising: The frequency domain information is converted back to the spatial domain through inverse discrete cosine transform, so that the image can be reconstructed into a visible form; The inverse discrete cosine transform formula is expressed as: ; in, represents the coefficients in the frequency domain, Represents the pixel value of the original image, x and y Respectively represent the horizontal and vertical coordinates of the image, and represents a frequency variable, N Indicates the size of the image. Represents the normalization factor.
Citation Information
Patent Citations
Underwater low-illumination image enhancement method based on conditional diffusion model
CN117911302A
Point cloud denoising method based on local DCT
CN118537256A