Image diffusion enhancement method and system based on decoupled guidance and reprojection refinement

By using a decoupling guidance and reprojection refinement method, and combining RGB and infrared thermal imaging images, a dual-condition temporal sensing U-Net network is constructed. This solves the problem of poor defogging effect in coal mine underground image defogging technology and achieves high-quality image diffusion enhancement.

CN121616477BActive Publication Date: 2026-03-31GUIZHOU INST OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-30
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

When existing image dehazing technology is applied in underground coal mines, it is difficult to balance dehazing effect, physical rationality, real-time performance and detail preservation, resulting in poor image diffusion enhancement effect.

Method used

By using a decoupled guidance and reprojection refinement method, a dual-condition temporal awareness U-Net network is constructed by fusing RGB images and infrared thermal images. Combined with a physical modulation module, a hierarchical cross-attention mechanism, and a reprojection refinement network, images are generated and corrected to enhance detail preservation and physical consistency.

Benefits of technology

It improves the stability and visual usability of image semantic extraction, generates natural images with fewer artifacts, conforms to the physical laws of atmospheric scattering, avoids smoothing effects, and improves image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121616477B_ABST
    Figure CN121616477B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image enhancement processing, in particular to an image diffusion enhancement method and system based on decoupling guidance and re-projection refinement. The method comprises: extracting mutually orthogonal physical degradation latent variables and semantic content latent variables through a two-way decoupling encoder; constructing a diffusion inverse process generation model containing a double conditional time sequence perception U-Net network, introducing physical consistency loss through a physical modulation module and a hierarchical cross-attention mechanism, generating a preliminary enhanced image, inputting the preliminary enhanced image into a re-projection refinement network for latent space residual correction; constructing a semantic-guided detail gain network to generate a detail gain map; performing detail enhancement through wavelet transform and the detail gain map to obtain a detail enhanced image output and generate a visual analysis report. The decoupling guidance ensures that the image generation process is strictly constrained, and the re-projection refinement corrects the deviation, enhances the image diffusion enhancement effect, and improves the image quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image enhancement processing technology, and in particular to an image diffusion enhancement method and system based on decoupling guidance and reprojection refinement. Background Technology

[0002] In the construction of modern and intelligent mines, visual perception systems serve as the "eyes" for monitoring the working face environment, identifying equipment status, analyzing personnel behavior, and enabling remote control and autonomous mining. However, the underground environment in coal mines is extremely harsh. The large amounts of coal dust generated during mining operations, the water mist stirred up by mechanical movement, and uneven tunnel lighting all contribute to a severe decline in the quality of video and images captured by cameras. Images typically exhibit low contrast, blurred details, and distorted colors, with valuable information buried in a "smog." This significantly limits the performance of subsequent advanced computer vision algorithms such as target detection and semantic segmentation, and even makes manual remote monitoring extremely difficult, constituting a key technological bottleneck in the construction of smart mines.

[0003] To address image degradation, existing image dehazing techniques have formed three main research directions: methods based on physical models, methods based on prior knowledge, and methods based on deep learning. Physical model-based methods rely on assumptions about the statistical characteristics of outdoor natural scenes. However, underground coal mine scenes have unique properties such as simple structure, complex distribution of artificial light sources, and strong non-uniformity of suspended particles (coal dust), making these assumptions difficult to apply. When applied to underground image dehazing, these methods often result in poor dehazing effects and image color distortion. While prior knowledge-based methods can improve image clarity to some extent, they suffer from inherent drawbacks such as high computational complexity and poor real-time performance, making them unsuitable for the real-time monitoring requirements of underground mining operations. Deep learning-based methods, especially those based on convolutional neural networks (CNNs) and generative adversarial networks (GANs), have made significant progress in image restoration tasks, but they still suffer from black-box problems and physical inconsistencies, pattern collapse and artifacts, and loss of detail and smoothing issues. To address this, the latest generative model—the Diffusion Model (DM)—has demonstrated superior generation quality and versatility compared to GANs in image generation tasks by simulating a forward noise-adding process from data to noise and a reverse noise-denoising process from noise to data. However, when using the diffusion model for image diffusion enhancement, it often simply uses a hazy image concatenated with noise as input, which does not provide the model with sufficiently refined and decoupled guiding information. Therefore, when performing image inpainting, it usually cannot preserve and enhance the original details to the maximum extent.

[0004] Therefore, traditional image diffusion enhancement methods are unable to adapt to the special and harsh environment of underground coal mines, and it is difficult to balance defogging effect, physical rationality, real-time performance and detail preservation. As a result, they suffer from poor image diffusion enhancement effect and low image quality. Summary of the Invention

[0005] In order to solve the above-mentioned technical problems, an image diffusion enhancement method and system based on decoupling guidance and reprojection refinement is provided. The decoupling guidance can ensure that the image generation process is strictly constrained, and the reprojection refinement can correct the deviation, enhance the image diffusion enhancement effect, and improve the image quality.

[0006] An image diffusion enhancement method based on decoupling guidance and reprojection refinement, the method comprising:

[0007] RGB images of dust-laden fog in the well and corresponding synchronous infrared thermal imaging images are acquired. The RGB images and infrared thermal imaging images are time-stamp aligned and spatially registered. Mutually orthogonal physical degradation latent variables and semantic content latent variables are extracted through a dual-channel decoupled encoder.

[0008] A diffusion inverse process generation model is constructed with a dual-condition temporal-aware U-Net network as the core for denoising. The physical degradation latent variable and semantic content latent variable are respectively input into the physical modulation module and hierarchical cross-attention mechanism of the diffusion inverse process generation model. Gradient guidance with physical consistency loss is introduced in the sampling stage to generate a preliminary enhanced image.

[0009] The preliminary enhanced image is input into the reprojection refining network of the encoding and decoding structure, latent space residual correction is performed through the residual correction network, and the refined image is output through the decoder.

[0010] A semantically guided detail gain network is constructed, and a detail gain map is generated based on the detail gain network. The refined image is then enhanced with detail through wavelet transform and the detail gain map to obtain a detail-enhanced image. The detail-enhanced image is then output and a visualization analysis report is generated.

[0011] In one embodiment, a dual-path decoupled encoder extracts mutually orthogonal physical degradation latent variables and semantic content latent variables, including:

[0012] A dual-channel decoupled encoder is constructed, and the foggy image after timestamp alignment and spatial registration is input into the dual-channel decoupled encoder;

[0013] The hazy image is decomposed into physical degradation latent variables and semantic content latent variables by the dual-path decoupling encoder, and a decoupling regularization loss based on Hilbert-Schmidt independence norm is introduced to ensure that the physical degradation latent variables and semantic content latent variables are spatially orthogonal.

[0014] In one embodiment, the decoupling regularization loss based on the Hilbert-Schmidt independence norm is: ;in, and Gram matrices representing the physical degradation latent variables and the semantic content latent variables, respectively; Represents a centered matrix; Represents the trace of a matrix.

[0015] In one embodiment, the method further includes:

[0016] The physical degradation latent variable is injected into the intermediate feature map of the network in the form of an affine transformation through the physical modulation module;

[0017] The semantic content latent variables are injected into the corresponding scale layer of the diffusion network decoder through the hierarchical cross-attention mechanism.

[0018] In one embodiment, the operation of the physical modulation module is defined as follows: ;in, and These are scaling factors and bias terms generated by the multilayer perceptron from the physical degradation latent variables; The intermediate feature map of the network is represented by C, H, and W, which are the number of channels, height, and width of the intermediate feature map of the network, respectively. This indicates element-wise multiplication by channel;

[0019] The calculation formula for the hierarchical cross-attention mechanism is as follows: ;in, Represents the first decoder in the dual-conditional time-aware U-Net network. Layer feature map; Indicates the first Latent variables of semantic content in layers; express Vector dimension; This indicates the matrix transpose.

[0020] In one embodiment, the physical consistency loss is defined as: ;in, In the noise reduction step Prediction results of intermediate clear images at that time; It is an atmospheric scattering physics model function. It is a parameter transmittance diagram. These are atmospheric light parameters. and From the decoder Decoded from the middle; It is structural similarity loss; It is an L1 or L2 pixel-level loss; It is a balancing weight; It is a foggy image after timestamp alignment and spatial registration.

[0021] In one embodiment, the training loss of the reprojection refinement network Including reconstruction loss Perceived loss Physical cycle consistency loss ,in: ; , All are loss weighting coefficients;

[0022] The operation definition for latent space residual correction is as follows: ;in, This indicates the initial image enhancement. Indicates via encoder The refined latent variables obtained; It is a physical degradation latent variable Semantic content latent variables As input, by parameters Optimized residual prediction network; Represents the corrected latent variables; the refined image is processed by the decoder. generate: .

[0023] In one embodiment, the detail gain map The update rule for adjusting high-frequency wavelet coefficients is as follows: ;in, For the first A high-frequency wavelet coefficient matrix of several scales; It is the first The detail gain map at each scale is generated by the detail gain network for the first... Each scale prediction is obtained; This indicates element-wise multiplication.

[0024] In one embodiment, the enhanced detail image is output and a visualization analysis report is generated, including:

[0025] Output the enhanced detail image, and generate a visualization analysis report based on the enhanced detail image, including image quality indicators, physical degradation distribution map, and enhanced detail region map;

[0026] The formula for generating the physical degradation distribution map is: ;in, is the scaling factor for the c-th channel; C is the total number of channels; This is an upsampling operation; These are the image space coordinates.

[0027] An image diffusion enhancement system based on decoupling guidance and reprojection refinement, the system comprising:

[0028] The data preprocessing and encoding module is used to acquire RGB images of dust and fog in the well and corresponding synchronous infrared thermal imaging images, align the RGB images and infrared thermal imaging images with timestamps and spatially register them, and extract mutually orthogonal physical degradation latent variables and semantic content latent variables through a dual-channel decoupled encoder.

[0029] The diffusion inverse process generation module is used to construct a diffusion inverse process generation model with a dual-condition temporal-aware U-Net network as the denoising core. The physical degradation latent variable and semantic content latent variable are respectively input into the physical modulation module and the hierarchical cross-attention mechanism of the diffusion inverse process generation model. Gradient guidance with physical consistency loss is introduced in the sampling stage to generate a preliminary enhanced image.

[0030] The reprojection refining module is used to input the preliminary enhanced image into the reprojection refining network of the encoding and decoding structure, perform latent space residual correction through the residual correction network, and output the refined image through the decoder.

[0031] The detail injection enhancement and analysis module is used to construct a semantically guided detail gain network and generate a detail gain map based on the detail gain network; the refined image is enhanced with detail through wavelet transform and the detail gain map to obtain a detail-enhanced image, and the detail-enhanced image is output and a visualization analysis report is generated.

[0032] The aforementioned image diffusion enhancement method and system based on decoupling guidance and reprojection refinement improves the stability of semantic extraction by fusing RGB and infrared information. The decoupled latent variables, representing degradation and content respectively, enhance the model's generalization ability and the independence of conditional control. Through a physical modulation module, hierarchical cross-attention, and online physical guidance, precise control of physical degradation and semantic content during generation is achieved, resulting in a natural, artifact-free, and atmosphericly scattering-compliant initially enhanced image. The reprojection refinement network further corrects potential physical deviations in diffusion generation. Detail enhancement through wavelet transform and detail gain maps achieves adaptive detail enhancement on top of dehazing, avoiding smoothing effects, improving image structural clarity and visual usability, enhancing image diffusion enhancement effects, and improving image quality. Attached Figure Description

[0033] Figure 1 This is a flowchart illustrating an image diffusion enhancement method based on decoupling guidance and reprojection refinement in one embodiment.

[0034] Figure 2 This is a schematic diagram of physical degradation visualization analysis in one embodiment;

[0035] Figure 3 This is a flowchart illustrating an image diffusion enhancement method based on decoupling guidance and reprojection refinement in another embodiment.

[0036] Figure 4 This is a schematic diagram showing the comparison of the image before and after enhancement in one embodiment;

[0037] Figure 5 This is a block diagram of an image diffusion enhancement system based on decoupling guidance and reprojection refinement in one embodiment.

[0038] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0040] The image diffusion enhancement method based on decoupling guidance and reprojection refinement provided in this application can be applied to computer equipment. The computer equipment can acquire RGB images of dust-laden areas and corresponding synchronous infrared thermal imaging images from underground. It performs time-stamp alignment and spatial registration between the RGB images and the infrared thermal imaging images, and extracts mutually orthogonal physical degradation latent variables and semantic content latent variables through a dual-path decoupling encoder. The computer equipment can construct a diffusion inverse process generation model with a dual-conditional temporal-aware U-Net network as the denoising core. The physical degradation latent variables and semantic content latent variables are input into the physical modulation module and hierarchical cross-attention mechanism of the diffusion inverse process generation model, respectively. Gradient guidance with physical consistency loss is introduced during the sampling stage to generate a preliminary enhanced image. The computer equipment can input the preliminary enhanced image into a reprojection refinement network with an encoder-decoder structure, perform latent spatial residual correction through a residual correction network, and output a refined image through a decoder. The computer equipment can construct a semantically guided detail gain network and generate a detail gain map based on the detail gain network. Detail enhancement is performed on the refined image through wavelet transform and the detail gain map to obtain a detail-enhanced image. The detail-enhanced image is then output and a visualization analysis report is generated. The computer equipment can include, but is not limited to, various personal computers, laptops, smartphones, robots, drones, tablets, and other similar devices.

[0041] In one embodiment, such as Figure 1 As shown, an image diffusion enhancement method based on decoupling guidance and reprojection refinement is provided, including the following steps:

[0042] Step 202: Acquire RGB images of dust-laden fog in the well and corresponding synchronous infrared thermal imaging images. Perform time stamp alignment and spatial registration between the RGB images and the infrared thermal imaging images. Extract mutually orthogonal physical degradation latent variables and semantic content latent variables through a dual-channel decoupled encoder.

[0043] In harsh environments such as coal mines and tunnels, camera systems composed of various camera devices can be deployed, connected to computer equipment, and transmitting the acquired image data to the computer for further processing. Specifically, the camera system includes at least an RGB camera and an infrared thermal imaging camera, both installed in the same location and with the same viewing angle.

[0044] The RGB camera is used to capture color visual information of the underground scene. Due to the influence of coal dust and water mist generated by mechanical agitation during underground mining operations, the captured images are RGB images containing dust and fog. The infrared thermal imaging camera captures infrared thermal images synchronized with the RGB images, which can effectively penetrate some of the dust and fog interference and retain stable structural information.

[0045] The computer device can record the acquisition timestamps of each frame of the RGB image and the infrared thermal imaging image, ensuring a one-to-one correspondence between the two in the time dimension. In this embodiment, the computer device can extract two sets of independent but crucial guidance information for image reconstruction from the acquired images: RGB images containing dust and water mist acquired by the downhole camera system, and optional synchronous infrared thermal imaging images. First, the computer device can receive the RGB foggy image acquired by the downhole camera. And optional, time-stamped and spatially registered in-scene infrared thermal images. Infrared images have better penetration through smoke and dust, providing more stable structural information for semantic coding. Next, computer equipment can perform time-stamp alignment and spatial registration between the RGB image and the infrared thermal image to obtain a foggy image after time-stamp alignment and spatial registration. .

[0046] In one embodiment, the provided image diffusion enhancement method based on decoupling guidance and reprojection refinement may further include a process of extracting latent variables. The specific process includes: constructing a dual-channel decoupling encoder and inputting the timestamp-aligned and spatially registered hazy image into the dual-channel decoupling encoder; decomposing the hazy image into physical degradation latent variables and semantic content latent variables through the dual-channel decoupling encoder, and introducing a decoupling regularization loss based on the Hilbert-Schmidt independence norm to ensure that the physical degradation latent variables and semantic content latent variables are spatially orthogonal.

[0047] To extract clean and non-interfering guidance information from a foggy image after timestamp alignment and spatial registration, a dual-channel decoupled encoder is constructed in this embodiment. The dual-channel decoupled encoder consists of two structurally independent and functionally specialized sub-encoders, which are used to extract physical degradation latent variables and hierarchical semantic content latent variables, respectively, and a specific constraint mechanism is used to ensure that the two types of latent variables are orthogonal to each other.

[0048] Specifically, a computer device can construct a dual-channel decoupled encoder, which consists of two parallel neural networks with different structures: a physical degradation encoder and a semantic content encoder. The physical degradation encoder... Employing a lightweight convolutional neural network structure, primarily composed of shallow convolutions and global pooling, the goal is to capture global features related to dust, fog, and lighting in images, such as overall color cast and contrast reduction, while ignoring specific scene content; a physical degradation encoder is used. Will Encoded as a low-dimensional physical degradation latent variable That is, the dual-channel decoupled encoder decomposes a foggy image into physical degradation latent variables representing fog concentration and light scattering, and semantic content latent variables representing scene content and structure. Among these, physical degradation visualization analysis includes... Figure 2 As shown, Figure 2 This indicates the distribution and severity of physical degradation (such as fog, dust, and vapor), with red / warm areas representing the most severe degradation and blue / cool areas representing milder degradation.

[0049] Semantic content encoder Employing a deep neural network structure, with a pre-trained Visual Transformer (ViT) or ConvNeXt as the backbone network, ensures powerful semantic feature extraction capabilities; semantic content encoder The main purpose is to extract robust structural and content information from the image that is not severely affected by fog. In other words, to enable subsequent hierarchical guidance, the encoder outputs a hierarchical semantic content latent variable. Each of them Feature maps corresponding to different depths (i.e. different scales) of the network.

[0050] In this embodiment, in order to ensure and True decoupling, to avoid interference between physically degraded information and semantic content information, and to ensure the purity of both types of latent variables, i.e. It only contains degradation information. It only contains content information. When training the encoder, in addition to the respective reconstruction or prediction tasks, a key decoupling regularization loss is introduced. That is, the decoupling regularization loss based on the Hilbert-Schmidt Independence Criterion (HSIC) is used to force physical degradation latent variables. Latent variables of hierarchical semantic content They are mutually orthogonal.

[0051] In one embodiment, the decoupling regularization loss based on the Hilbert-Schmidt independence norm is: ;in, and Gram matrices representing latent variables of physical degradation and semantic content, respectively; Represents a centered matrix; Represents the trace of a matrix.

[0052] in, and These are physical and semantic latent variables extracted from a batch of samples of size N. and Matrix; and Based on Gaussian kernel function The calculated Gram matrix, i.e. Calculated using the Gaussian kernel function; It is a centralized matrix. It is the identity matrix. It is an all-1 vector. The decoupling regularization loss term based on the Hilbert-Schmidt independence norm forces the encoder to decouple by minimizing the cross-correlation of latent variables in the Hilbert space of the regeneration kernel. In this embodiment, this is achieved by adding and minimizing the total loss function. It can effectively punish and Any statistical correlation between them forces the two encoders to learn uncorrelated feature representations.

[0053] Step 204: Construct a diffusion inverse process generation model with a dual-condition temporal-aware U-Net network as the core of denoising. Input the physical degradation latent variable and the semantic content latent variable into the physical modulation module and the hierarchical cross-attention mechanism of the diffusion inverse process generation model, respectively. In the sampling stage, introduce gradient guidance with physical consistency loss to generate a preliminary enhanced image.

[0054] Computer devices can employ a diffusion model as a generator, generating a dual-condition temporal-aware U-Net network (DCTA-U-Net) based on a dual-condition injection and guided diffusion inverse process. This network serves as the denoising core of the diffusion model, which includes a fixed forward process and a learnable backward process. In the forward process, noise is gradually added to the clear image until it reaches a completely random state using a fixed Gaussian noise scheduling strategy. The backward process, relying on the DCTA-U-Net network, gradually removes noise from pure noise under dual-condition guidance of physical degradation latent variables and hierarchical semantic content latent variables, restoring the clear image structure and details.

[0055] Specifically, in the forward process, Within each time step, gradually moving towards a clearer image. Adding Gaussian noise yields a series of noisy images. ; at any time Noisy images It can be done through formula Directly obtained, among which It's noise. These are predefined noise scheduling coefficients. The reverse process is the part the model needs to learn. During the reverse process, a dual-conditional time-aware U-Net (DCTA-U-Net) can be constructed at each time step. According to the noisy image Physical degradation latent variables The network is injected with physical modulation (Phys-Mod) modules, and hierarchical semantic latent variables are added. Noise added to the image is predicted by injecting through a hierarchical cross-attention mechanism. .

[0056] Meanwhile, in this embodiment, an online guiding term is also introduced during the sampling process. The gradient of the guiding term comes from a dynamic physical consistency loss, which forces the intermediately generated clear image to approximate the original foggy image after physical degradation is reapplied, ensuring that the generation process does not deviate from physical reality.

[0057] In one embodiment, the provided image diffusion enhancement method based on decoupling guidance and reprojection refinement may further include a dual-condition injection process, specifically including: physical degradation latent variables being injected into the intermediate feature map of the network through a physical modulation module in the form of an affine transformation; and semantic content latent variables being injected into the corresponding scale layer of the diffusion network decoder through a hierarchical cross-attention mechanism.

[0058] Computer devices can detect global physical degradation latent variables. The hierarchical semantic latent variables are injected into each residual block of the U-Net through a specially designed Phys-Mod module. The corresponding scale layer of the U-Net decoder (upsampling path) is injected through a hierarchical cross-attention mechanism.

[0059] In one embodiment, the operation of the physical modulation module is defined as follows: ;in, and These are scaling factors and bias terms generated by the multilayer perceptron from the physical degradation latent variables; This represents the intermediate feature map of the network, where C, H, and W represent the number of channels, height, and width of the intermediate feature map, respectively. This represents element-wise multiplication across channels; the calculation formula for the hierarchical cross-attention mechanism is: ;in, Represents the first decoder in a dual-conditional time-aware U-Net network. Layer feature map; Indicates the first Latent variables of semantic content in layers; express Vector dimension; This indicates the matrix transpose.

[0060] In DCTA-U-Net, the Phys-Mod module is similar to AdaIN in style transfer. It contains two small multilayer perceptrons (MLPs) whose function is to convert global physical degradation latent variables. Transform into feature map The parameters for performing the affine transformation will be Mapped to a scaling factor equal to the number of channels in the feature map of that block. and bias terms : This allows the model to dynamically adjust the statistical properties of each layer of features based on the overall fog concentration, thereby achieving global defogging intensity control.

[0061] The calculation method of the hierarchical cross-attention mechanism is as follows: in the U-Net decoder... Each scale layer, with its upsampled feature map As a query, and from The semantic vectors extracted at the corresponding scale After linear transformation, these become the Key and Value, which are then calculated. The model can focus on and fuse the most relevant semantic information at specific spatial locations and scales in the generated image, thereby ensuring that the generated content has a correct structure and clear outline.

[0062] In addition to conditional injection, an online guidance mechanism is introduced during the inference sampling phase to enhance physical consistency. This mechanism is implemented at each sampling step. The model first predicts the noise. This allows for the calculation of the current best estimate for the final sharp image. Then calculate a physical consistency loss. The gradient is used to "push" the current value. Then proceed with the next sampling step.

[0063] In one embodiment, the physical consistency loss is defined as: ;in, In the noise reduction step Prediction results of intermediate clear images at that time; It is an atmospheric scattering physics model function. It is a parameter transmittance diagram. These are atmospheric light parameters. and From the decoder Decoded from the middle; It is the structural similarity (SSIM) loss; It is an L1 or L2 pixel-level loss; It is a balancing weight; This is a hazy image after timestamp alignment and spatial registration. Physical consistency loss measures "how similar a hazy image is to the original input if it were resynthesized using the currently predicted sharp image and the predicted fog parameters." It is corrected by moving along the direction of the fastest decrease in physical consistency loss (the negative gradient direction). This ensures that the entire denoising process is always "anchored" to physical reality. Through iterative steps, a preliminary, clear image was finally obtained. In each sampling step, the physical consistency loss is calculated based on... Find the gradient To correct the sampling direction.

[0064] Step 206: The preliminarily enhanced image is input into the reprojection refining network of the encoding and decoding structure, latent space residual correction is performed through the residual correction network, and the refined image is output through the decoder.

[0065] Despite the generated initial enhanced image The quality is already very high, but the randomness of the diffusion process may still introduce some tiny artifacts that do not conform to physical laws. To address this, a Stochastic Reprojection Refinement Network (SRR-Net) was designed to correct these deviations. The reprojection refinement network is an autoencoder with an encoder-decoder structure.

[0066] Computer equipment can refine the initially enhanced image based on the physical consistency of random latent space reprojection. The input is fed into a random reprojection refinement network (SRR-Net); the reprojection refinement network first passes through an encoder Will Mapping to a compact, refined latent space yields latent variables. Then, using a latent variable derived from physical degradation... and semantic content latent variables Commonly parameterized correction network , Correction network take over and the original physical latent variables and semantic latent variables As input, the refined latent variables are subjected to residual correction, and a residual vector is output. .

[0067] In one embodiment, the operation of latent space residual correction is defined as follows: ;in, This indicates the initial image enhancement. Indicates via encoder The refined latent variables obtained; It is a physical degradation latent variable Semantic content latent variables As input, by parameters Optimized residual prediction network; Represents the corrected latent variables; the refined image is processed by the decoder. generate: That is, latent space residual correction utilizes the most original and purest conditional information ( Proofreading by Encoded, and may contain noise. Finally, through the decoder The revised Reconstructed into a refined image with higher physical fidelity. .

[0068] In one embodiment, the training loss of the reprojected refined network is... Including reconstruction loss Perceived loss Physical cycle consistency loss ,in: ; , All of these are loss weighting coefficients.

[0069] The training of the Stochastic Reprojection Refined Network (SRR-Net) is performed using a composite loss function. Supervision, training loss It consists of three parts: It is a refined image With real and clear images L1 reconstruction loss between; The perceptual loss is calculated in the feature space of a pre-trained VGG network to ensure consistency in visual perception. It is the physical cycle consistency loss, defined as: The training loss of the constructed reprojection refined network. Forced Refined Image After "re-fogging" through the physical model, it must be compared with the original input. The high degree of consistency ensures that the SRR-Net correction process is not arbitrary, but strictly serves the ultimate goal of improving physical fidelity.

[0070] Step 208: Construct a semantically guided detail gain network and generate a detail gain map based on the detail gain network; enhance the details of the refined image by wavelet transform and detail gain map to obtain a detail-enhanced image, output the detail-enhanced image and generate a visualization analysis report.

[0071] Computer devices can inject and enhance semantic details into the output image based on controllable wavelet transform, thus refining the image. A multi-level discrete wavelet transform (DWT) is performed to decompose the image into low-frequency approximate components and multi-scale high-frequency detail components. A detail injection network is then constructed based on the original foggy image. High-frequency components and semantic latent variables As input, a detail gain map is predicted, which is then used for adaptive augmentation. The high-frequency detail components are extracted; finally, the final detail-enhanced image is reconstructed through inverse discrete wavelet transform (IDWT). .

[0072] Specifically, since fog and detail enhancement are inherently contradictory, in order to achieve detail enhancement while dehazing, this embodiment designs a post-processing module that operates in the frequency domain. This post-processing module can perform wavelet decomposition, detail gain prediction, detail injection, and wavelet reconstruction. Wavelet decomposition is primarily used for refining the image. Performing a multi-level discrete wavelet transform (DWT) yields a low-frequency approximate component. And a series of high-frequency detail components of different scales and directions (horizontal, vertical, diagonal) Detail gain prediction mainly involves constructing a lightweight detail injection network. The input to the detail injection network is the original foggy image. Corresponding high-frequency wavelet coefficients and semantic latent variables The output is a multi-scale detail gain map. Semantic information The introduction of this feature enables the network to identify which regions have important details that should be enhanced; detail injection primarily uses the predicted gain map for adaptive adjustment. High-frequency coefficients; wavelet reconstruction mainly uses the updated high-frequency coefficients. and the original low-frequency coefficient The final detail-enhanced image is reconstructed using inverse discrete wavelet transform (IDWT). .

[0073] In one embodiment, detail gain map The update rule for adjusting high-frequency wavelet coefficients is as follows: ;in, For the first High-frequency wavelet coefficient matrix at various scales (e.g., horizontal, vertical, diagonal); It is the first The detail gain map at each scale is generated by the detail gain network for the first... Each scale prediction is obtained; This represents element-wise multiplication. By adjusting the high-frequency wavelet coefficients using a detail gain map, selective detail enhancement can be achieved in key regions of an image determined by semantic information (such as equipment outlines and rock textures). In regions with high values, high-frequency components are amplified, thus enhancing details; in regions with low values, details remain unchanged or are even slightly suppressed to remove noise.

[0074] In one embodiment, the provided image diffusion enhancement method based on decoupling guidance and reprojection refinement may further include a process of outputting a detail-enhanced image and generating an analysis report. The specific process includes: outputting the detail-enhanced image; and generating a visualization analysis report based on the detail-enhanced image, including image quality indicators, a physical degradation distribution map, and a detail-enhanced region map. The formula for generating the physical degradation distribution map is: ;in, is the scaling factor for the c-th channel; C is the total number of channels; This is an upsampling operation; These are the image space coordinates.

[0075] Among them, the physical degradation distribution map The generation method is as follows: the scaling factor of all physical modulation (Phys-Mod) modules in DCTA-U-Net is used. The L2 norm is aggregated along the channel dimension and upsampled to the original image size. The physical degradation distribution map intuitively reflects the intensity of the model's dehazing operation in different regions of the image, with high-value regions corresponding to areas perceived as dense fog by the model.

[0076] Specifically, the aforementioned image diffusion enhancement method based on decoupling guidance and reprojection refinement can be integrated into a complete, deployable end-to-end processing system. The system can process real-time video streams or offline image files from downhole cameras; that is, it receives real-time video streams or offline images, outputs enhanced clear images, and calculates image quality metrics (such as UIQM and NIQE) before and after enhancement. Simultaneously, the system can calculate objective quality evaluation metrics for the images before and after enhancement, such as reference-free UIQM (Underwater Image Quality Measurement) and NIQE (Natural Image Quality Evaluator) scores. In this embodiment, the system can utilize intermediate information within the model to generate a visual analysis map, which may include a physical degradation distribution map and a detail enhancement region map. The physical degradation distribution map... The generation process is as follows: by aggregating the activation strengths (such as scaling factors) of all Phys-Mod modules in DCTA-U-Net. The norm of the algorithm can be used to generate a heatmap that visually shows which areas in the original image the model considers to have the densest fog. (Detail enhancement region map) During the generation process: directly use the predicted detail gain map The visualization reveals which areas the model has focused on enhancing in terms of detail. This information, along with the enhanced images, is integrated into an interactive HTML report and presented as the final visualization analysis report.

[0077] In one embodiment, the overall flow of an image diffusion enhancement method based on decoupling guidance and reprojection refinement is as follows: Figure 3 As shown, the specific process includes:

[0078] RGB images of dust-laden fog in the well and corresponding synchronous infrared thermal images were acquired and time-stamped and spatially registered.

[0079] Multi-source input and dual-channel decoupled coding: Multi-source images are input into a dual-channel decoupled encoder to extract mutually orthogonal physical degradation latent variables and semantic content latent variables;

[0080] Biconditionally guided diffusion inverse process generation: A diffusion inverse process generation model with a biconditionally temporally aware U-Net network as the core of denoising is constructed. The physical degradation latent variable and the semantic content latent variable are respectively input into the physical modulation module and the hierarchical cross attention mechanism of the diffusion inverse process generation model. Gradient guidance of physical consistency loss is introduced in the sampling stage to generate a preliminary enhanced image.

[0081] Latent space reprojection refinement: The initially enhanced image is input into the reprojection refinement network of the encoding and decoding structure, latent space residuals are corrected through the residual correction network, and the refined image is output through the decoder;

[0082] Wavelet domain semantic detail injection: Construct a semantically guided detail gain network and generate a detail gain map based on the detail gain network; enhance the details of the refined image by wavelet transform and detail gain map to obtain a detail-enhanced image;

[0083] System integration and report generation: Output groups of images with enhanced details and generate visual analysis reports.

[0084] In one embodiment, a demonstration process is provided for single-frame image processing using an image diffusion enhancement method based on decoupled guidance and reprojection refinement, with the input being a 640x480 image filled with coal dust captured by a tunneling machine camera. The specific processing mainly includes several parts: encoding, diffusion generation, image refining, image enhancement, and report output.

[0085] The encoding process includes: It is fed into the dual-channel decoupled encoder, the physical encoder. Output a 128-dimensional... Semantic encoder Output semantic vectors from three layers (corresponding to different scales of U-Net). Meanwhile, HSIC loss guarantee and Irrelevant.

[0086] The diffusion process includes: from a random Gaussian noise image Begin with 100 steps of back diffusion; in the... Step (such as) DCTA-U-Net receiver Time encoding, and and ; The activation values ​​of each layer of the network are adjusted using Phys-Mod, and Cross-attention is applied to the three layers of the decoder respectively; network prediction noise. Then, calculate the online guided gradient. right Make corrections, then sample to obtain Repeat this process to obtain a preliminary clear image. The generated, pre-refined image is used as the enhanced image, and the unprocessed original image is used as the unenhanced image. The comparison between the two is as follows: Figure 4 As shown, the enhanced image is significantly of higher quality.

[0087] The image refining process includes: Encoded as 256 dimensions by SRR-Net ; Correction network utilization Predict a correction amount ,get ; decoder from generate .

[0088] The image enhancement process includes: [The text abruptly ends here, likely due to an incomplete sentence or a formatting error.] Perform a 3-level wavelet transform; inject detail network according to High-frequency information and Predicting a three-layer gain plot ;Will Take a ride The final image is obtained by applying inverse wavelet transform to the high-frequency coefficients. .

[0089] The report output process includes: output It displays the UIQM score as increasing from 0.8 (input) to 2.5 (output); it also generates two heatmaps: one showing the highest physical degradation at the top and right sides of the image (in the direction of the dust source); and the other showing the strongest detail gain at the outline of the tunneling machine cutter and hydraulic support.

[0090] In this embodiment, an image diffusion enhancement method based on decoupling guidance and reprojection refinement is provided. By fusing RGB and infrared information, the stability of semantic extraction can be improved. The decoupled latent variables represent degradation and content respectively, which can enhance the model's generalization ability and the independence of conditional control. Through physical modulation module, hierarchical cross attention, and online physical guidance, precise control of physical degradation and semantic content in the generation process can be achieved. The generated initial enhanced image is natural, has few artifacts, and conforms to the physical laws of atmospheric scattering. The reprojection refinement network can further correct the physical deviations that may exist in diffusion generation. Detail enhancement is performed through wavelet transform and detail gain map, achieving adaptive enhancement of details on the basis of dehazing, avoiding smoothing effects, improving the structural clarity and visual usability of the image, enhancing the image diffusion enhancement effect, and improving image quality.

[0091] This application presents an image diffusion enhancement method based on decoupling guidance and reprojection refinement. Employing a diffusion model as the generation backbone, its generation capability surpasses that of GANs, producing more natural images with fewer artifacts. More importantly, through physical-semantic decoupling, physical modulation injection, and online physical consistency guidance, the generation process is strictly constrained by physical laws at every step. The final reprojection refinement step further corrects physical deviations, resulting in a recovered image that is not only clear but also highly consistent with physical reality in terms of illumination and transmittance. This application abandons the traditional paradigm that inevitably leads to smoothing of details during dehazing. Through semantic decoupling and hierarchical cross-attention, the model can understand and preserve important structures in the image. The unique wavelet transform-based detail injection step can decompose high-frequency information from the original blurred image and selectively and controllably enhance it according to semantic guidance, making key details such as device edges and rock textures sharper and clearer after dehazing. The physical-semantic decoupling design allows the model to cope with changes in content and fog conditions. Instead of learning a rigid mapping from specific fog to a specific scene, the model learns the general physical inverse process of "defogging" itself, and how to adjust it according to the scene content. This gives the application strong generalization ability and robustness to dust and fog environments of different concentrations and types, as well as diverse underground scenes. This application completely breaks away from the GAN framework, avoiding its problems such as training instability and mode collapse, resulting in a more stable training process and better convergence of the diffusion model. Furthermore, through visualization outputs such as physical degradation distribution maps, users can clearly see which regions the model performs stronger defogging operations, enhancing the model's credibility and practicality.

[0092] It should be understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0093] In one embodiment, such as Figure 5 As shown, an image diffusion enhancement system based on decoupling guidance and reprojection refinement is provided, including: a data preprocessing and encoding module 610, a diffusion inverse process generation module 620, a reprojection refinement module 630, and a detail injection enhancement and analysis module 640, wherein:

[0094] The data preprocessing and encoding module 610 is used to acquire RGB images of dust and fog in the well and corresponding synchronous infrared thermal imaging images, align the RGB images and infrared thermal imaging images with timestamps and spatially register them, and extract mutually orthogonal physical degradation latent variables and semantic content latent variables through a dual-channel decoupled encoder.

[0095] The diffusion inverse process generation module 620 is used to construct a diffusion inverse process generation model with a dual-condition temporal-aware U-Net network as the denoising core. The physical degradation latent variable and semantic content latent variable are respectively input into the physical modulation module and the hierarchical cross-attention mechanism of the diffusion inverse process generation model. Gradient guidance of physical consistency loss is introduced in the sampling stage to generate a preliminary enhanced image.

[0096] The reprojection refining module 630 is used to input the initially enhanced image into the reprojection refining network of the encoding and decoding structure, perform latent space residual correction through the residual correction network, and output the refined image through the decoder.

[0097] The detail injection enhancement and analysis module 640 is used to construct a semantically guided detail gain network and generate a detail gain map based on the detail gain network; the refined image is enhanced with detail through wavelet transform and detail gain map to obtain a detail-enhanced image, and the detail-enhanced image is output and a visualization analysis report is generated.

[0098] In one embodiment, the data preprocessing and encoding module 610 is further used to construct a dual-channel decoupled encoder and input the timestamp-aligned and spatially registered foggy image into the dual-channel decoupled encoder; the dual-channel decoupled encoder decomposes the foggy image into physical degradation latent variables and semantic content latent variables, and introduces a decoupling regularization loss based on the Hilbert-Schmidt independence norm to ensure that the physical degradation latent variables and semantic content latent variables are spatially orthogonal.

[0099] In one embodiment, the diffusion inverse process generation module 620 is further used to inject physical degradation latent variables into the intermediate feature map of the network in the form of affine transformation through the physical modulation module; and to inject semantic content latent variables into the corresponding scale layer of the diffusion network decoder through a hierarchical cross-attention mechanism.

[0100] In one embodiment, the detail injection enhancement and analysis module 640 is further configured to output a detail-enhanced image and generate a visualization analysis report based on the detail-enhanced image, including image quality indicators, physical degradation distribution map, and detail-enhanced region map.

[0101] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 6As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When executed by the processor, the computer program implements an image diffusion enhancement method based on decoupled booting and reprojection refinement. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0102] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0103] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of an image diffusion enhancement method based on decoupled guidance and reprojection refinement.

[0104] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of an image diffusion enhancement method based on decoupled guidance and reprojection refinement.

[0105] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0106] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0107] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An image diffusion enhancement method based on decoupled guidance and reprojection refinement, characterized in that, The method comprises: Collecting the RGB image and the corresponding synchronous infrared thermal imaging image of the dust-containing fog downhole, time stamping and spatially registering the RGB image and the infrared thermal imaging image, and extracting mutually orthogonal physical degradation latent variables and semantic content latent variables through a dual-decoupling encoder, introducing a decoupling regularization loss based on the Hilbert-Schmidt independence norm, and ensuring that the physical degradation latent variables and the semantic content latent variables are spatially orthogonal; The diffusion inverse process generation model with the dual-condition time sequence perception U-Net network as the denoising core is constructed, and the physical degradation latent variable and the semantic content latent variable are respectively input into a physical modulation module and a hierarchical cross-attention mechanism of the diffusion inverse process generation model, including that the physical degradation latent variable is injected into the intermediate feature map of the network in the form of affine transformation through the physical modulation module, and the semantic content latent variable is injected into the corresponding scale layer of the diffusion network decoder through the hierarchical cross-attention mechanism; in the sampling stage, gradient guidance of a physical consistency loss is introduced to generate a preliminary enhanced image; the physical consistency loss is defined as: ; wherein, indicates an intermediate clear image prediction result in a denoising step . is an atmospheric scattering physical model function, is a parameter transmittance map, is an atmospheric light parameter, and are obtained by decoding the decoder from . is a structural similarity loss; is an L1 or L2 pixel-level loss; is a balance weight; is a foggy image after time stamp alignment and spatial registration. inputting the preliminary enhanced image into a reprojection refining network of a codec structure, correcting a latent space residual through a residual correction network, and outputting a refined image through a decoder; a training loss of the reprojection refining network including a reconstruction loss , a perception loss , a physical cycle consistency loss , wherein ; , are loss weight coefficients; an operation definition of the latent space residual correction is ; wherein represents a preliminary enhanced image, represents a refined latent variable obtained through an encoder ; is a residual prediction network optimized by parameters , taking a physical degradation latent variable and a semantic content latent variable as inputs; represents a corrected latent variable; the refined image is generated through a decoder : ; A semantically guided detail gain network is constructed, and a detail gain map is generated based on the detail gain network. The refined image is then enhanced with details using wavelet transform and the detail gain map to obtain a detail-enhanced image. The detail-enhanced image is output, and a visualization analysis report is generated. The detail gain map... The update rule for adjusting high-frequency wavelet coefficients is as follows: ;in, For the first A high-frequency wavelet coefficient matrix of several scales; It is the first The detail gain map at each scale is generated by the detail gain network for the first... Each scale prediction is obtained; This indicates element-wise multiplication.

2. The image diffusion enhancement method based on decoupled guidance and reprojection refinement of claim 1, wherein, Extracting mutually orthogonal physical degradation latent variables and semantic content latent variables through a dual-decoupling encoder, comprising: Building a dual-decoupling encoder, and inputting the foggy image after time stamping and spatially registering into the dual-decoupling encoder; Decomposing the foggy image into physical degradation latent variables and semantic content latent variables through the dual-decoupling encoder.

3. The image diffusion enhancement method based on decoupled guidance and reprojection refinement of claim 2, wherein, The decoupling regularization loss based on the Hilbert-Schmidt independence norm is: ; wherein, and respectively represent the Gram matrices of the physical degradation latent variable and the semantic content latent variable; represents a centering matrix; represents a trace of a matrix.

4. The image diffusion enhancement method based on decoupled guidance and reprojection refinement of claim 1, wherein, The operation of the physical modulation module is defined as: ; wherein, and are scaling factors and bias terms generated by a multi-layer perceptron from the physical degradation latent variable; represents the network intermediate feature map, C, H, and W are the channel number, height, and width of the network intermediate feature map, respectively; represents element-by-element multiplication of channels. The calculation formula of the hierarchical cross-attention mechanism is: ; wherein, represents the first layer feature map of the decoder in the dual-condition time sequence perception U-Net network; represents the semantic content latent variable of the first layer; represents vector dimension; represents matrix transposition.

5. The image diffusion enhancement method based on decoupled guidance and reprojection refinement of claim 1, wherein, Outputting the detail-enhanced image and generating a visual analysis report, comprising: Outputting the detail-enhanced image, and generating a visual analysis report containing image quality indicators, physical degradation distribution maps, and detail-enhanced region maps based on the detail-enhanced image; The generation formula of the physical degradation distribution map is: ; wherein, is a scaling factor of the cth channel; C is the total number of channels; is an up-sampling operation; is an image space coordinate.

6. An image diffusion enhancement system based on decoupled guidance and reprojection refinement for implementing the method of image diffusion enhancement based on decoupled guidance and reprojection refinement according to any one of claims 1 to 5, characterized in that, The system comprises: A data preprocessing and encoding module for collecting the RGB image and the corresponding synchronous infrared thermal imaging image of the dust-containing fog downhole, time stamping and spatially registering the RGB image and the infrared thermal imaging image, and extracting mutually orthogonal physical degradation latent variables and semantic content latent variables through a dual-decoupling encoder; A diffusion inverse process generation module for building a diffusion inverse process generation model with a dual-condition time sequence perception U-Net network as the denoising core, inputting the physical degradation latent variables and the semantic content latent variables into a physical modulation module and a hierarchical cross-attention mechanism of the diffusion inverse process generation model, respectively, and introducing gradient guidance of physical consistency loss in the sampling stage to generate a preliminary enhanced image; A re-projection refining module for inputting the preliminary enhanced image into a re-projection refining network of a coding and decoding structure, performing latent space residual correction through a residual correction network, and outputting a refined image through a decoder; A detail injection enhancement and analysis module for building a semantic-guided detail gain network, generating a detail gain map based on the detail gain network, performing detail enhancement on the refined image through wavelet transform and the detail gain map to obtain a detail-enhanced image, outputting the detail-enhanced image, and generating a visual analysis report.

Citation Information

Patent Citations

  • Extreme weather photovoltaic power prediction method, system, equipment and medium

    CN120805092A

  • Image super-resolution method and system based on semantic perception token

    CN121353083A