Mine fully mechanized coal mining face low-illumination image zero reference enhancement method and system

By combining the diffusion model reverse denoising and the Prediction-Net network, the image enhancement is achieved using the illumination edge invariant and the content invariant, which solves the problem of unsatisfactory image enhancement effect in the low-illumination environment of the mine comprehensive mining working surface, and effectively enhances and recognizes low-light images.

CN120147162AActive Publication Date: 2025-06-13CHINA UNIV OF MINING & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510224865.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In the low illumination environment of the mine comprehensive mining surface, traditional low-light image enhancement methods are difficult to effectively deal with complex lighting conditions, serious noise and loss of image details, resulting in unsatisfactory enhancement effect.

Method used

The reverse denoising method of diffusion model is used to enhance image by using the Prediction-Net network through information guidance during prediction and denoising process, combining the illumination edge invariant and the content invariant. This method does not rely on labeled data, and only uses normal optical image data sets for training, which has good generalization capabilities.

Benefits of technology

It realizes effective enhancement of low-light images, adapts to complex and changeable lighting conditions, improves the retention and recognition accuracy of image details, and is suitable for image enhancement in low-illumination environments of the mine comprehensive mining working surface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147162A_ABST
    Figure CN120147162A_ABST
Patent Text Reader

Abstract

The invention discloses a zero reference enhancement method and system for a low-illumination image of a fully mechanized coal mining face of a mine, and the method comprises the steps: starting from a noise map through a reverse denoising part of a pre-training diffusion model, and achieving the guidance of the denoising process content through a prediction estimation value; during training, only a normal light image data set is used, and illumination edge invariants and content invariants of low-light and normal light images and an estimated value x0 of the tth step are extracted to serve as input of the network, so that the network can fully learn normal illumination features during training and restore a source image according to input information. During testing, the network completes construction of illumination features according to invariant information and continuously refines an estimated value x0, and low-light image enhancement is achieved. The system comprises a multispectral image acquisition module, an illumination invariant extraction module, a diffusion model initialization and noise generation module, a Prediction-Net network processing module, a reverse denoising guide module and a post-processing optimization module. According to the invention, the low-light image can be effectively enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a zero-reference enhancement method and system for low-illumination images in fully-mechanized coal mining faces in mines, belonging to the technical field of low-light image enhancement. Background Art

[0002] With the continuous increase in the depth of coal mining, the underground environment has become increasingly complex. Especially in fully-mechanized coal mining faces, in the case of poor lighting conditions, the image recognition technology relied on by coal mining machinery and personnel operations faces severe challenges. In low-illumination environments, due to insufficient light, a large amount of image details are missing in conventional vision systems, affecting the recognition of mining equipment, personnel, and the mine environment, and further hindering the precise control and scheduling of automated mining machinery. Studying the low-illumination image enhancement technology for fully-mechanized coal mining faces in coal mines is of great significance for improving the level of coal mine safety production, increasing mining efficiency, and ensuring the safety of miners' lives.

[0003] At present, traditional low-light image enhancement methods have many limitations. For example, histogram equalization improves contrast by redistributing image gray values, but it is easy to over-enhance certain areas, resulting in the loss of image details and the appearance of block effects, and it cannot specifically process the complex low-illumination and severely noisy images of fully-mechanized coal mining faces in mines, and the enhancement effect is not satisfactory; although the gamma transformation method can improve image brightness and contrast, its adaptability to uneven illumination is poor and it is difficult to meet the complex and changeable lighting conditions underground in mines; in the field of deep learning, supervised methods require a large number of paired low-light and normal-light images for training. However, it is difficult to obtain such data, and the enhancement results depend on the training data. Once there are differences between the training data and the actual mine image distribution, the generalization ability of the model will be reduced. Unsupervised methods generally use GAN models. Although they do not require labeled data, the training is difficult, the adversarial balance is difficult to achieve, and problems such as gradient disappearance or explosion are likely to occur. In addition, whether it is traditional methods or deep learning-based methods, it is difficult to handle complex image noise, reflected light, and object occlusion in the low-illumination environment of mines. Summary of the Invention

[0004] The invention purpose of the present invention is to provide a zero-reference enhancement method and system for low-illumination images in fully-mechanized coal mining faces in mines. This method and system can effectively enhance low-light images, can adapt to complex and changeable lighting conditions, and have strong generalization ability.

[0005] To achieve the above purpose, the present invention provides a zero-reference enhancement method for low-illumination images in fully-mechanized coal mining faces in mines, including the following steps:

[0006] S1. Diffusion model reverse denoising: Starting from a noise map that conforms to the standard normal distribution, predict the enhanced result conditional on x t at the t-th step of denoising Obtain the denoising result of the (t - 1)-th step, and use this information to guide the image generation of the (t - 1)-th step, so as to obtain the final generated result;

[0007] S2. Model training: Input the normal light image, and extract the illumination edge invariant and content invariant;

[0008] S3. Input the illumination edge invariant and the x t estimated by x 0|t into the Prediction-Net network to predict Minimize and the MSE loss of the normal light image, and by analogy, obtain x t-1 , …, x 1 , and predict

[0009] S4. During testing, input the low-light image, repeat S2 and S3 to obtain the final enhancement result.

[0010] Furthermore, in the denoising process in S1, x 0|t is estimated from the denoising result of the t-th step, and the formula is as follows:

[0011]

[0012] where α t = 1 - x t , x t is a predefined value, which is a parameter controlling the noise addition intensity in each diffusion process, with a magnitude in the range of (0, 1), and gradually decreases as t increases, ∈ θ (x t , t) is the noise predicted at the t-th step, and x t represents the result obtained at the (t + 1)-th step;

[0013] From obtain the denoising result of the (t - 1)-th step, and use this information to guide the image generation of the (t - 1)-th step, and the formula is as follows:

[0014]

[0015] where N represents the Gaussian distribution, I is the noise map randomly sampled from the standard normal distribution at the t-th step, and are the mean and variance of the intermediate noise map at the t-th step respectively, and the formulas are as follows:

[0016]

[0017] The x t-1 obtained through the above steps is consistent with the input image in content, and at the same time, through xt-1 Obtain x t-2 , and so on, and finally obtain the enhanced result.

[0018] Furthermore, the illumination edge invariant described in S2 is derived according to the Kubelka-Munk theory. Assuming equal energy but non-uniform illumination and a dull and matte object surface, the invariant N is derived. At the same time, assuming equal energy and uniform illumination and a dull and matte object surface, the invariant W is derived. The formulas are as follows:

[0019]

[0020] where E x and E y are the horizontal and vertical gradients of E respectively; E λx and E λy are the horizontal and vertical gradients of E λ respectively; E λλx and E λλy are the horizontal and vertical gradients of E λλ respectively; E λ and E λλ are the first and second derivatives of E with respect to the wavelength λ respectively,

[0021] E, E λ , E λλ are calculated by the following formulas respectively:

[0022]

[0023] where R(x, y), G(x, y), B(x, y) are the gray values at the corresponding positions of the three channels of the original image;

[0024] The illumination content invariant is the frequency map obtained after the original image is Fourier-transformed. The formula is:

[0025] A I = |FFT(input)|;

[0026] where FFT represents the Fourier transform.

[0027] Combining the illumination edge invariant and the content invariant as the input of the network can better maintain the structural and content relevance between low-light and normal-light images.

[0028] Furthermore, the Prediction-Net network in S3 adopts the U-Net architecture, which includes three downsamplings and three upsamplings, and uses skip connections to fuse features at different levels. It is expressed by the formula as follows:

[0029] ds = DownSample(x1 );

[0030] us = UpSample(x 2 );

[0031] f_result = Conv(Concat(ds, us));

[0032] Wherein, DownSample represents the downsampling operation, UpSample represents the upsampling operation, Concat is the concatenation operation, and Conv is the convolution operation;

[0033] The U-Net architecture further includes an APFM module, and the APFM module adopts a multi-stage fusion method to refine and modify the phase diagrams of A I and x 0|t while performing inverse Fourier transform on the respective corresponding phase diagrams and amplitude diagrams, and finally fusing the feature diagrams obtained in each stage, expressed as:

[0034] f c = Relu(Conv(x)) * 3;

[0035] fussion 1 = IFFT(f c ((|FFT(x 0|t )|), A I ));

[0036] fussion 2 = IFFT(|FFT(x 0|t )|, f c (A I ));

[0037] fussion 3 = IFFT(|FFT(x 0|t )|, A I ));

[0038] fussion r = Relu(Conv(Concat(fussion 1 , fussion w , fussion 3 )));

[0039] Wherein, Relu is the activation function, *3 means the previous operation is repeated 3 times, and IFFT is the inverse Fourier transform;

[0040] During training, minimize the mean square error between and the normal light image, and the formula is:

[0041]

[0042] Among them, represents the final result of the t-th step prediction. I(i, j) represents the input normal light image, and M and N represent the width and height of the image respectively;

[0043] The exposure control loss is used to control the deviation of the gray intensity and the middle tone value, and the formula is:

[0044]

[0045] Among them, Y k represents the average gray value of the local area, P represents the number of region blocks, and E is a constant;

[0046] The total loss is:

[0047] L = MSE + α * L e .

[0048] The overall function of the network is mainly to predict the original normal light image according to the given illumination edge invariant and content invariant of the input and the x estimated at the t-th step 0|t , and predict the original normal light image. This method is a zero-reference method, which is trained only using the normal light image dataset and has good generalization ability.

[0049] Furthermore, in S4, by extracting the illumination edge invariant and content invariant of the low-light image, and at the same time using the illumination prior information learned by the Prediction-Net network, after the low-light image is input into the network, through continuous denoising and prediction the final enhancement result is obtained.

[0050] A zero-reference enhancement system for low-illumination images in a fully-mechanized coal mining face includes a multi-spectral image acquisition module, an illumination invariant extraction module, a diffusion model initialization and noise generation module, a Prediction-Net network processing module, a reverse denoising guidance module, and a post-processing optimization module;

[0051] The multi-spectral image acquisition module is deployed at key positions of the shearer and hydraulic supports, and synchronously acquires the original RAW image stream through a visible light low-illumination camera and a near-infrared auxiliary light source array, and transmits it to the illumination invariant extraction module through a multi-sensor synchronous controller;

[0052] The illumination invariant extraction module uses the Kubelka-Munk theory to calculate the illumination edge invariants N and W, and combines the Fourier transform to extract the content invariant A I , generates a fused feature map and inputs it into the Prediction-Net module; when the system starts, the diffusion model initialization module generates a standard normal distribution noise map x T, and adjust the reverse denoising process through time-step iteration;

[0053] The described Prediction-Net network processing module adopts a U-Net architecture, and fuses the illumination invariant and the diffusion intermediate result x through the APFM module 0|t , and predicts the enhanced image after frequency-domain phase correction FFT / IFFT and feedback it to the reverse denoising guidance module; this module iteratively updates the Gaussian distribution parameter μ t , generate the x t-1 sequence until the preliminary enhanced result x 0 is output;

[0054] The described post-processing optimization module compresses the dynamic range through differentiable Tone Mapping, performs color correction by combining SSIM and color constancy loss, and adjusts the enhancement intensity based on the real-time feedback of the mine dust concentration, and finally outputs the result to the monitoring terminal.

[0055] In the present invention, by using the reverse denoising part of the pre-trained diffusion model starting from the noise map, the denoising process is guided in terms of content through predicted estimated values. At the same time, in order to better adapt to the complex illumination environment of the fully mechanized coal mining face in the mine, only the normal light image dataset is used during training. By extracting the illumination edge invariant, content invariant, and the estimated value x at the t-th step of the low-light and normal-light images 0 as the input of the network, the network can fully learn the normal illumination features during training and restore the source image according to the input information. During testing, the network can complete the construction of the illumination features based on the invariant information and continuously refine the estimated value x 0 , and finally realize the enhancement of low-light images. The present invention effectively enhances low-light images, adapts to complex and variable illumination conditions, has strong generalization ability, and can be widely applied to image enhancement in low-illumination environments of fully mechanized coal mining faces in mines. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is a schematic diagram of the working process of the method of the present invention;

[0057] Figure 2 is a framework diagram of the method of the present invention;

[0058] Figure 3 is a framework diagram of the Prediction-Net network in the embodiment of the present invention;

[0059] Figure 4 is a diagram of the APFM module in the embodiment of the present invention;

[0060] Figure 5 is a comparison diagram of the low-light image enhancement effect in the embodiment of the present invention. Detailed implementation mode

[0061] The present invention will be further described below with reference to the accompanying drawings.

[0062] As Figure 1 and Figure 2 shown, a zero-reference enhancement method for low-illumination images of fully-mechanized mining faces in mines includes the following steps:

[0063] S1. Diffusion model reverse denoising: Starting from a noise map that conforms to the standard normal distribution, at the t-th step of denoising, predict the enhanced result t conditioned on x to obtain the denoising result of the (t - 1)-th step, and use this information to guide the image generation of the (t - 1)-th step, so as to obtain the final generation result;

[0064] S2. Model training: Input a normal light image, and extract the illumination edge invariant and content invariant;

[0065] S3. Input the illumination edge invariant and the x t estimated by x 0|t into the Prediction-Net network, predict minimize the MSE loss with the normal light image, and so on to obtain x t-1 , …, x 1 , predict

[0066] S4. During testing, input a low-light image, repeat S2 and S3 to obtain the final enhanced result.

[0067] Furthermore, in the denoising process of S1, x 0|t is estimated from the denoising result of the t-th step, and the formula is as follows:

[0068]

[0069] Wherein, α t = 1 - β t β t is a predefined value, ∈ θ (x t , t) is the noise predicted at the t-th step, and x t represents the result obtained at the (t + 1)-th step;

[0070] From obtain the denoising result of the (t - 1)-th step, and use this information to guide the image generation of the (t - 1)-th step, and the formula is as follows:

[0071]

[0072] Among them, N represents the Gaussian distribution, and I is the noise map randomly sampled from the standard normal distribution at the t-th step. and are the mean and variance of the intermediate noise map at the t-th step respectively, and the formulas are as follows:

[0073]

[0074] The x obtained through the above steps t-1 is consistent with the input image in content. At the same time, through x y-1 we get x t-2 , and so on, and finally the enhancement result is obtained.

[0075] Furthermore, as Figure 3 shown, the illumination edge invariant described in S2 is derived according to the Kubelka-Munk theory. Assuming equal energy but non-uniform illumination and a dull and matte object surface, the invariant N is derived. At the same time, assuming equal energy and uniform illumination and a dull and matte object surface, the invariant W is derived. The formulas are as follows:

[0076]

[0077] Among them, E x and E y are the horizontal and vertical gradients of E respectively; E λx and E λy are the horizontal and vertical gradients of E λ respectively; E λλx and E λλy are the horizontal and vertical gradients of E λλ respectively; E λ and E λλ are the first and second derivatives of E with respect to the wavelength λ respectively.

[0078] E, E λ , E λλ are calculated by the following formulas respectively:

[0079]

[0080] Among them, R(x, y), G(x, y), and B(x, y) are the gray values at the corresponding positions of the three channels of the original image respectively;

[0081] The illumination content invariant is the frequency map obtained after the original image undergoes Fourier transform. The formula is:

[0082] A I = |FFT(iuput)|;

[0083] Among them, FFT represents the Fourier transform;

[0084] Using the comprehensive illumination edge invariant and content invariant as the input of the network can better maintain the structural and content correlation between low-light and normal-light images.

[0085] Furthermore, as Figure 4 shown, the Prediction-Net network in S3 adopts a U-Net architecture, which includes three downsamplings and three upsamplings, and uses skip connections to fuse features at different levels, which is expressed by the formula as follows:

[0086] ds = DownSample(x 1 );

[0087] us = UpSample(x 2 );

[0088] f_result = Conv(Concat(ds, us));

[0089] Among them, DownSample represents the downsampling operation, UpSample represents the upsampling operation, Concat is the concatenation operation, and Conv is the convolution operation;

[0090] The U-Net architecture also includes an APFM module. The APFM module uses a multi-stage fusion method to refine and modify the phase diagrams of A I and x 0|t , and at the same time performs inverse Fourier transforms on the corresponding phase diagrams and amplitude diagrams, and finally fuses the feature diagrams obtained in each stage, which is expressed as:

[0091] f c = Relu(Conv(x)) * 3;

[0092] fussion 1 = IFFT(f c ((|FFT(x 0|t )|), A I );

[0093] fussion 2 = IFFT(|FFT(x 0|t )|, f c (A I ));

[0094] fussion 3 = IFFT(|FFT(x 0|t )|, A I );

[0095] fussion r= Relu(Conv(Concat(fussion 1 , fussion 2 , fussion 3 )));

[0096] Among them, Relu is the activation function, *3 means the previous operation is repeated 3 times, and IFFT is the inverse Fourier transform;

[0097] During training, minimize the mean square error between

[0098]

[0099] and the normal light image. The formula is: where

[0100] represents the final result of the prediction at the t-th step, I(i, j) represents the input normal light image, and M and N respectively represent the width and height of the image;

[0101]

[0102] where Y k represents the average gray value of a local area of size 16×16 in this embodiment, P represents the number of region blocks, and E is a constant, which is taken as 0.6 in the present invention.

[0103] The overall function of the network is mainly to predict the original normal light image according to the given illumination edge invariant and content invariant of the input and x estimated at the t-th step 0|t . This method is a zero-reference method, which is trained only using the normal light image dataset and has good generalization ability.

[0104] Furthermore, in S4, by extracting the illumination edge invariant and content invariant of the low-light image, and simultaneously using the illumination prior information learned by the Prediction-Net network, after the low-light image is input into the network, through continuous denoising and prediction the final enhancement result is obtained.

[0105] The present invention also provides a zero-reference enhancement system for low-illumination images of fully-mechanized coal mining faces in mines, including a multi-spectral image acquisition module, an illumination invariant extraction module, a diffusion model initialization and noise generation module, a Prediction-Net network processing module, a reverse denoising guidance module, and a post-processing optimization module;

[0106] The described multi - spectral image acquisition module is deployed at key positions of shearers and hydraulic supports. It synchronously acquires the original RAW image stream through a visible - light low - illumination camera and a near - infrared auxiliary light source array, and transmits it to the illumination invariant extraction module through a multi - sensor synchronous controller;

[0107] The described illumination invariant extraction module uses the Kubelka - Munk theory to calculate the illumination edge invariants N and W, and combines Fourier transform to extract the content invariant A I , generates a fused feature map and inputs it into the Prediction - Net module; When the system starts, the diffusion model initialization module generates a standard normal distribution noise map x T , and adjusts the reverse denoising process through time - step iteration;

[0108] The described Prediction - Net network processing module adopts a U - Net architecture, fuses the illumination invariant and the diffusion intermediate result x through the APFM module 0|t , predicts the enhanced image after frequency - domain phase correction FFT / IFFT and feeds it back to the reverse denoising guidance module; This module iteratively updates the Gaussian distribution parameter μ t , generates the x t-1 sequence until the preliminary enhanced result x is output 0 ;

[0109] The described post - processing optimization module compresses the dynamic range through differentiable Tone Mapping, performs color correction by combining SSIM and color constancy loss, and adjusts the enhancement intensity in real - time based on the mine dust concentration feedback, and finally outputs the result to the monitoring terminal.

[0110] This embodiment is trained on the COCO - 2017 dataset, and low - light images on the LOL dataset are selected as test data. The results are as Figure 5 shown: Three types of low - light images are selected, namely indoor images, outdoor images, and mine low - light images. Among them, Figure 5 (a), (b) represent the indoor low - light image and the enhanced result respectively; Figure 5 (c), (d) represent the outdoor low - light image and the enhanced result respectively; Figure 5 (e), (f) are the low - light image and the enhanced result under a real mine respectively. It can be seen that the present invention has achieved significant enhancement of low - light images in multiple scenarios.

Claims

1. A zero-reference enhancement method for low-illumination images of fully-mechanized mining working faces in mines, characterized in that: The steps include: S1, diffusion model reverse denoising: starting from the noise map that conforms to the standard normal distribution, predict x in the tth step of denoising t Conditional enhancement results Get the denoising result of step t-1, use this information to guide the image generation of step t-1, and then get the final generation result; S2, model training: input normal light image, extract illumination edge invariant and content invariant; S3, the illumination edge invariant and the t The estimated x 0|t Input Prediction-Net network, prediction minimize and the MSE loss of the normal light image, and so on to get x t-1 ,…,x1, prediction S4: During testing, input a low-light image and repeat S2 and S3 to obtain the final enhanced result.

2. The method for zero-reference enhancement of low-illumination images of fully mechanized mining working faces in mines according to claim 1 is characterized in that: In the denoising process of S1, x 0|t It is estimated from the denoising result of step t, and the formula is as follows: in, α t =1-β t , β t is a predefined value, ∈ θ (x t ,t) is the noise predicted in the tth step, x t represents the result obtained in step t+1; Depend on The denoising result of step t-1 is obtained, and this information is used to guide the image generation of step t-1. The formula is as follows: Where N represents Gaussian distribution, I is the noise image randomly sampled from the standard normal distribution at step t, and are the mean and variance of the intermediate noise map in the tth step, respectively, and the formulas are: The x obtained by the above steps is t-1 The content is consistent with the input image, and through x t-1 Get x t-2 , and finally get the enhanced result.

3. The method for zero-reference enhancement of low-illumination images of fully mechanized mining working faces in a mine according to claim 1 is characterized in that: The illumination edge invariant described in S2 is derived according to the KubelkaMunk theory. Assuming equal energy but uneven illumination and dull surface of the object, the invariant N is derived. At the same time, assuming equal energy and uniform illumination and dull surface of the object, the invariant W is derived. The formulas are: Among them, E x and E y are the horizontal and vertical gradients of E, respectively; E λx and E λy E λ The horizontal and vertical gradients of E λλx and E λλy E λλ The horizontal and vertical gradients of E λ and E λλ are the first and second derivatives of E with respect to wavelength λ, E, E λ , E λλ They are calculated by the following formulas: Among them, R(x,y), G(x,y), and B(x,y) are the grayscale values ​​of the corresponding positions of the three channels of the original image respectively; The illumination content invariant is the frequency map obtained after Fourier transform of the original image. The formula is: A I =|FFT(input)|; Here, FFT stands for Fourier transform.

4. The method for zero-reference enhancement of low-illumination images of fully mechanized mining working faces in a mine according to claim 3 is characterized in that: The Prediction-Net network in S3 adopts the U-Net architecture, which includes three downsampling and three upsampling, and uses skip connections to fuse features at different levels, which can be expressed as follows: ds = DownSample(x1); us = UpSample(x2); f_result=Conv(Concat(ds,us)); Among them, DownSample represents the downsampling operation, UpSample represents the upsampling operation, Concat represents the concatenation operation, and Conv represents the convolution operation; The U-Net architecture also includes an APFM module, which uses a multi-stage fusion method to refine and modify A I and x 0|t The phase map is obtained by performing inverse Fourier transform with the corresponding phase map and amplitude map, and finally the feature maps obtained at each stage are fused, which is expressed as: f c =Relu(Conv(x))*3; fussion1=IFFT(f c ((|FFT(x 0|t )|),A I ); fussion2=IFFI(|FFT(x 0|t )|,f c (A I )); fussion3=IFFT(|FFT(x 0|t )|,A I ); merger r =Relu(Conv(Concat(fusion1,fusion2,fusion3))); Among them, Relu is the activation function, *3 means the previous operation is repeated 3 times, and IFFT is the inverse Fourier transform; Minimize during training The mean square error between the image and the normal light image is: in, represents the final result of the prediction at step t, I(i,j) represents the input normal light image, M and N represent the width and height of the image respectively; Exposure control loss is used to control the deviation of grayscale intensity and mid-tone value. The formula is: Among them, Y k represents the average gray value of the local area, P represents the number of regional blocks, and E is a constant; The total loss is: L=MSE+α*L e 。 5. The method for zero-reference enhancement of low-illumination images of fully mechanized mining working faces in a mine according to claim 1, characterized in that: In S4, the illumination edge invariant and content invariant of the low-light image are extracted, and the illumination prior information learned by the Prediction-Net network is used to input the low-light image into the network, and then the low-light image is continuously denoised and predicted. Thus the final enhanced result is obtained.

6. A zero-reference enhancement system for low-light images of fully-mechanized mining working faces in mines, characterized in that: It includes multispectral image acquisition module, illumination invariant extraction module, diffusion model initialization and noise generation module, Prediction-Net network processing module, reverse denoising guidance module and post-processing optimization module; The multispectral image acquisition module is deployed at key positions of coal mining machines and hydraulic supports, and the original RAW image stream is synchronously acquired through a visible light low-light camera and a near-infrared auxiliary light source array, and transmitted to the illumination invariant extraction module through a multi-sensor synchronization controller; The illumination invariant extraction module uses the Kubelka-Munk theory to calculate the illumination edge invariants N and W, and combines Fourier transform to extract the content invariant A I , generate a fusion feature map and input it into the Prediction-Net module; when the system starts, the diffusion model initialization module generates a standard normal distribution noise map x T , and adjust the reverse denoising process through time step iteration; The Prediction-Net network processing module adopts the U-Net architecture and fuses the illumination invariant and the diffusion intermediate result x through the APFM module. 0|t , after frequency domain phase correction FFT / IFFT, the predicted enhanced image And feed back to the reverse denoising guidance module; this module iteratively updates the Gaussian distribution parameter μ t , Generate x t-1 sequence until the initial enhancement result x0 is output; The post-processing optimization module compresses the dynamic range through differentiable Tone Mapping, performs color correction by combining SSIM with color constancy loss, and adjusts the enhancement intensity based on real-time feedback of mine dust concentration, and finally outputs the results to the monitoring terminal.

Citation Information

Patent Citations

  • Low-light image enhancement method based on zero-reference Retinex decomposition network

    CN119168895A

  • Low-light image enhancement method based on deep Retinex

    JP7493867B1

  • Apparatus and method for low-light image enhancement with generative adversarial network based denoising function

    KR102611606B1

  • Low-light image enhancement method based on reinforcement learning and aesthetic evaluation

    WO2023236565A1