A tunnel water mist environment under inspection image de-scattering enhancement and reconstruction method

By combining neural networks and frequency domain attention networks, and utilizing polarization characteristics and inspection pose information, image descattering enhancement and reconstruction under tunnel water mist conditions are achieved, solving the problems of tunnel image blurring and color distortion, and improving image clarity and robustness.

CN122115250APending Publication Date: 2026-05-29CHINA RAILWAY LIUYUAN GRP CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA RAILWAY LIUYUAN GRP CO LTD
Filing Date
2026-04-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing image processing methods in tunnel water mist environments suffer from problems such as image blurring, reduced contrast, and color distortion. Furthermore, deep learning methods lack generalization ability in real tunnel environments and exhibit poor robustness when processing dynamic water mist scenarios.

Method used

A fusion neural network is used to decouple the scene, and an adaptive weight is generated by combining a frequency domain attention network. Through multi-scale frequency domain fusion and visual optimization processing, static scene branches and dynamic water mist polarization branches are constructed using polarization characteristics and inspection pose information to generate high-quality enhanced and reconstructed images.

Benefits of technology

It effectively removes the effects of water mist in tunnel inspection images, improves image clarity and detail, and produces more realistic results with fewer artifacts, stronger robustness, and adaptability to complex and ever-changing water mist environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115250A_ABST
    Figure CN122115250A_ABST
Patent Text Reader

Abstract

The application discloses a tunnel water mist environment under inspection image de-scattering enhancement and reconstruction method, and belongs to the technical field of image processing. The method comprises the following steps: acquiring a polarized image sequence and pose information collected synchronously; decoupling a scene by using a fusion neural network to output a fog-free image, a geometric gradient map and a water mist concentration map; calculating a frequency domain fusion weight map by using a frequency domain attention generation network; performing multi-scale frequency domain decomposition on the fog-free image to obtain high-frequency components and low-frequency components; performing pixel-level weighted fusion and inverse transformation according to the weight map to generate a preliminary enhanced image and perform color correction. The application adopts a physical constraint deep learning and a frequency domain adaptive fusion mechanism, can effectively eliminate the interference of tunnel water mist scattering, and improves the definition, detail performance and color authenticity of the inspection image, thereby providing high-quality visual support for tunnel safety monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment. Background Technology

[0002] Image processing refers to the technology of using computers to analyze, process, and manipulate images to meet specific visual needs. In many application scenarios, such as tunnel inspection, the quality of acquired images directly affects subsequent analysis and decision-making. Due to its enclosed and humid environment, tunnels often contain water mist. Light scatters in the mist, causing images captured by inspection robots or vehicles to appear blurry, have reduced contrast, and exhibit color distortion, severely impacting the accurate assessment of the tunnel lining structure and ancillary facilities. Therefore, developing image descattering enhancement and reconstruction technologies specifically for tunnel water mist environments is of great significance for ensuring the safety of tunnel operation and maintenance.

[0003] Existing image dehazing techniques are mainly divided into two categories. One category is image enhancement-based methods, such as histogram equalization and the Retinex algorithm. These methods directly adjust the pixel values ​​of the image to improve the visual effect. Although computationally simple, they usually do not consider the physical model of image degradation, which can easily lead to loss of image details, color distortion, or noise. The other category is physical model-based methods, such as dark channel prior and polarization dehazing. These methods estimate the transmittance map and atmospheric light by establishing an atmospheric scattering model, and then reconstruct a clear image. Although these methods can restore image quality to a certain extent, they often rely on strong prior assumptions. In atypical scenarios such as tunnels with complex artificial lighting and uneven water mist distribution, these assumptions are difficult to meet, resulting in unstable dehazing effects.

[0004] In recent years, with the development of deep learning technology, end-to-end image dehazing methods based on convolutional neural networks have made progress. These methods can automatically extract features and complete image restoration by learning the mapping relationship between a large number of foggy and fog-free image pairs. However, most existing deep learning methods rely on synthetic datasets for training, and synthetic data differs significantly from the physical characteristics of real tunnel water mist environments, resulting in insufficient model generalization ability. In addition, most methods treat dehazing as a single image-to-image conversion task, failing to fully utilize the geometric information of the scene or the physical characteristics of water mist scattering. The processing results often lack physical interpretability and have poor robustness when dealing with dynamically changing water mist scenes. Summary of the Invention

[0005] To address the aforementioned issues, this invention provides a method for descattering enhancement and reconstruction of inspection images in tunnel water mist environments. It employs a fusion neural network for scene decoupling, a frequency domain attention network to generate adaptive weights, and combines multi-scale frequency domain fusion with visual optimization processing. This method effectively removes the influence of water mist in tunnel inspection images, improving image clarity and detail.

[0006] To achieve the above objectives, this application adopts the following technical solution: Firstly, a method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment is provided, comprising: acquiring a sequence of images with polarization characteristics synchronously collected along the tunnel inspection path and corresponding inspection pose information; constructing a fusion neural network with a static scene branch and a dynamic water mist polarization branch, inputting the image sequence with polarization characteristics and the inspection pose information into the fusion neural network for scene information decoupling, and outputting a fog-free scene image, a geometric gradient map reflecting the scene structural features, and a water mist concentration distribution map reflecting the degree of environmental degradation; and constructing a frequency domain attention generation network for extracting scene features. The geometric gradient map and the water mist concentration distribution map are input into a frequency domain attention generation network. Based on the scene geometry and water mist distribution pattern, a spatially adaptive frequency domain fusion weight map is calculated. The fog-free scene image is decomposed into a multi-scale frequency domain to obtain high-frequency components representing image details and low-frequency components representing image structure and brightness. Based on the frequency domain fusion weight map, the high-frequency components and the low-frequency components are weighted and fused at the pixel level, and an inverse transform is performed to generate a preliminary enhanced image. The preliminary enhanced image is then subjected to color correction and contrast equalization to output the final enhanced and reconstructed image.

[0007] Based on the above technical solution, in the method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment provided in this application, a fusion neural network is used for scene decoupling, a frequency domain attention network is used to generate adaptive weights, and multi-scale frequency domain fusion and visual optimization processing are combined to effectively remove the influence of water mist in tunnel inspection images and improve image clarity and detail.

[0008] In conjunction with the first aspect described above, in one possible implementation, the method further includes: performing an inter-frame consistency check on the acquired image sequence with polarization characteristics to filter out valid image frames; using the acquired inspection pose information to perform smoothing filtering on the pose corresponding to the valid image frames to generate smoothed inspection trajectory information; and using the image sequence composed of the valid image frames and the smoothed inspection trajectory information as input data for subsequent steps to improve scene decoupling and enhance the stability of reconstruction.

[0009] In conjunction with the first aspect above, in one possible implementation, the construction of the fusion neural network with a static scene branch and a dynamic water mist polarization branch includes: using a tunnel image degradation simulator containing a differentiable physical scattering model to generate a simulated polarization image sequence and a corresponding simulated water mist concentration map based on a clear tunnel image; inputting the simulated polarization image sequence into the fusion neural network to be trained, minimizing the image reconstruction loss between the fog-free scene image and the clear tunnel image, and minimizing the concentration estimation loss between the output water mist concentration distribution map and the simulated water mist concentration map through a backpropagation algorithm, thereby completing the network parameter update.

[0010] In conjunction with the first aspect mentioned above, in one possible implementation, inputting the image sequence with polarization characteristics and the inspection pose information into a fusion neural network for scene information decoupling includes: performing Stokes vector calculation on the image sequence with polarization characteristics collected on-site to obtain a polarization degree map reflecting the intensity of water mist scattering and a polarization angle map reflecting the vibration direction of polarized light; using the polarization degree map, polarization angle map, and inspection pose information as joint input data for the trained fusion neural network; performing three-dimensional implicit modeling of the tunnel scene based on the inspection pose information through the static scene branch, rendering and outputting a fog-free scene image, and deriving the scene geometric gradient map by differentiating the implicit function; and learning the physical modulation model of water mist on polarized light through the dynamic water mist polarization branch, and combining the polarization degree map and polarization angle map for physical consistency constraints, and inverting and outputting a water mist concentration distribution map.

[0011] In conjunction with the first aspect above, in one possible implementation, the construction of the frequency domain attention generation network for extracting scene features includes: constructing an operator for edge enhancement processing of the scene's geometric gradient map, and a mapping unit for nonlinear mapping processing of the water mist concentration distribution map; establishing a multi-channel stitching layer at the output of the operator and the mapping unit, and connecting several convolutional layers after the multi-channel stitching layer to fuse the detail preservation weight map generated by the operator with the fog suppression weight map generated by the mapping unit to extract a composite feature map reflecting the scene structure and the water mist distribution pattern; setting an attention mechanism layer containing two parallel convolutional branches after the convolutional layers, wherein the two parallel convolutional branches are respectively configured to calculate the high-frequency attention score and low-frequency attention score at each pixel position based on the composite feature map, forming a spatially adaptive frequency domain fusion weight map; using the simulated water mist concentration map generated by the tunnel image degradation simulator as training data, and updating the parameters of the frequency domain attention generation network by calculating the correlation loss between the weight distribution output by the frequency domain attention generation network and the simulated water mist concentration map.

[0012] In conjunction with the first aspect above, in one possible implementation, the differentiable physical scattering model includes: a scattering coefficient related to ambient humidity and a polarization attenuation coefficient related to water mist particle size; during the execution of the backpropagation algorithm, the values ​​of the physical parameters are dynamically adjusted by calculating the gradient of the total loss relative to the scattering coefficient and the polarization attenuation coefficient; by updating the physical parameters, the feature distribution of the simulated polarization image sequence is adjusted in a direction that makes the simulated degradation closer to the distribution of real acquired data, thereby enhancing the generalization ability of the fused neural network and the frequency domain attention generation network.

[0013] In conjunction with the first aspect described above, in one possible implementation, performing pixel-level weighted fusion of the high-frequency components and the low-frequency components, and then performing inverse transform processing to generate a preliminary enhanced image includes: extracting high-frequency channel weights corresponding to the high-frequency components and low-frequency channel weights corresponding to the low-frequency components from the spatially adaptive frequency domain fusion weight map output by the frequency domain attention generation network; multiplying the high-frequency channel weights with the high-frequency components pixel by pixel to obtain a weighted high-frequency sub-band image, thereby preserving details in edge regions and suppressing noise in flat regions; multiplying the low-frequency channel weights with the low-frequency components pixel by pixel to obtain a weighted low-frequency sub-band image, thereby restoring illumination in foggy regions and avoiding overexposure in well-lit regions; adding the weighted high-frequency sub-band image and the weighted low-frequency sub-band image pixel by pixel to obtain a fused frequency domain sub-band image; and performing a multi-scale frequency domain inverse transform on the fused frequency domain sub-band image to recombine the frequency domain features and convert them back to the spatial domain to generate a preliminary enhanced image.

[0014] In conjunction with the first aspect mentioned above, in one possible implementation, performing color correction and contrast equalization processing on the preliminary enhanced image to output the final enhanced reconstructed image includes: converting the preliminary enhanced image from the RGB color space to the HSV color space and extracting the luminance component map therein; performing contrast-limited adaptive histogram equalization processing on the luminance component map to enhance the local contrast of the image while suppressing noise amplification, generating an equalized luminance component; merging the equalized luminance component with the original hue and saturation components in the HSV color space, and converting the merged data back to the RGB color space to output the final enhanced reconstructed image.

[0015] In conjunction with the first aspect mentioned above, in one possible implementation, the dynamic water mist polarization branch is used to learn the physical modulation model of water mist on polarized light, and the polarization degree map and polarization angle map are combined for physical consistency constraints. The inversion output of the water mist concentration distribution map includes: sampling the three-dimensional space within the current camera's view frustum based on the inspection pose information and the scene surface distance information generated by the static scene branch; estimating the water mist density at each sampling point and the polarization modulation parameters characterizing the modulation effect of water mist particles on the polarization state of light; using the estimated water mist density and polarization modulation parameters at each sampling point, combined with a fog-free scene image, simulating the change in polarization characteristics of light passing through the water mist through a volume rendering equation, and rendering a predicted polarization degree map and a predicted polarization angle map corresponding to the current viewpoint; calculating the difference between the predicted polarization degree map and the actually calculated polarization degree map, and the difference between the predicted polarization angle map and the actually calculated polarization angle map, to obtain the polarization physical loss; updating the weights of the dynamic water mist polarization branch using the polarization physical loss, and integrating the optimized three-dimensional water mist density field along the line of sight to generate a two-dimensional water mist concentration distribution map.

[0016] Secondly, a system for image descattering enhancement and reconstruction in tunnel water mist environment is provided, comprising: a synchronous acquisition module for acquiring image sequences with polarization characteristics and corresponding inspection pose information synchronously acquired along the tunnel inspection path; a scene information decoupling module for constructing a fusion neural network with static scene branches and dynamic water mist polarization branches, inputting the image sequences with polarization characteristics and the inspection pose information into the fusion neural network for scene information decoupling, and outputting a fog-free scene image, a geometric gradient map reflecting the scene structure features, and a water mist concentration distribution map reflecting the degree of environmental degradation; and a frequency domain weight calculation module for constructing a frequency domain attention generation network for extracting scene features. The geometric gradient map and the water mist concentration distribution map are input into the frequency domain attention generation network. Based on the scene geometry and water mist distribution pattern, a spatially adaptive frequency domain fusion weight map is calculated. A multi-scale frequency domain decomposition module is used to perform multi-scale frequency domain decomposition on the fog-free scene image to obtain high-frequency components representing image details and low-frequency components representing image structure and brightness. A pixel-level weighted fusion module is used to perform pixel-level weighted fusion of the high-frequency and low-frequency components based on the frequency domain fusion weight map, and perform inverse transformation processing to generate a preliminary enhanced image. A visual quality optimization module is used to perform color correction and contrast equalization processing on the preliminary enhanced image, outputting the final enhanced reconstructed image.

[0017] Compared with the prior art, the present invention has the following advantages: This invention constructs a fusion neural network with parallel static and dynamic branches, utilizing polarized image sequences and inspection pose information to effectively decouple tunnel scenes, water mist distribution, and scene geometric features. This decoupling method, based on multimodal data and physical model constraints, fundamentally separates the inherent attributes of the scene from the instantaneous influence of the environment, providing high-quality intermediate data for precise enhancement. Compared to traditional methods, its processing results are more realistic, with fewer artifacts and stronger robustness.

[0018] This invention proposes a frequency-domain attention generation network that dynamically calculates spatially adaptive frequency-domain fusion weights based on the decoupled scene geometric gradient and water mist concentration map. This mechanism enables the image enhancement process to intelligently differentiate between different regions of the image, focusing on restoring brightness and structure in areas with dense water mist, and on preserving texture and edges in areas rich in detail. This achieves refined and differentiated processing of high-frequency details and low-frequency structures in the image, effectively avoiding common problems in traditional enhancement methods such as detail blurring, noise amplification, or overexposure.

[0019] This invention introduces a training strategy based on a differentiable physical scattering model, which dynamically optimizes the physical parameters in the scattering model during training. This allows the simulated training data to continuously approximate the degradation characteristics of real water mist environments, thereby improving the generalization ability of the neural network model to real, complex, and variable water mist environments. This adaptive optimization mechanism, which combines data generation and model training, effectively solves the problem of obtaining large amounts of paired training data in real-world scenarios, improving the practicality and reliability of the entire method.

[0020] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A structural architecture diagram of an inspection image descattering enhancement and reconstruction system in a tunnel water mist environment provided in this application embodiment; Figure 2 A flowchart illustrating a method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment, provided in an embodiment of this application; Figure 3 This is a heatmap of spatial adaptive frequency domain fusion weight distribution provided in the embodiments of this application.

[0023] Figure 4 This is a training convergence curve of the differentiable physical parameters provided in the embodiments of this application. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] The method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment provided in this application embodiment can be applied to, for example... Figure 1 In the tunnel water mist environment inspection image descattering enhancement and reconstruction system 100 shown, such as Figure 1 As shown, the system includes: Synchronous acquisition module: used to acquire image sequences with polarization characteristics and corresponding inspection pose information synchronously acquired in the tunnel inspection path; Scene information decoupling module: used to construct a fusion neural network with static scene branch and dynamic water mist polarization branch, input the image sequence with polarization characteristics and the inspection pose information into the fusion neural network to decouple the scene information, and output fog-free scene image, geometric gradient map reflecting scene structural features and water mist concentration distribution map reflecting the degree of environmental degradation; Frequency domain weight calculation module: used to construct a frequency domain attention generation network for extracting scene features. The geometric gradient map and the water mist concentration distribution map are input into the frequency domain attention generation network. Based on the scene geometry and water mist distribution pattern, a spatially adaptive frequency domain fusion weight map is calculated. Multi-scale frequency domain decomposition module: used to perform multi-scale frequency domain decomposition on the fog-free scene image to obtain high-frequency components representing image details and low-frequency components representing image structure and brightness; Pixel-level weighted fusion module: used to perform pixel-level weighted fusion of the high-frequency components and the low-frequency components according to the frequency domain fusion weight map, and perform inverse transformation processing to generate a preliminary enhanced image; Visual quality optimization module: used to perform color correction and contrast equalization on the preliminary enhanced image, and output the final enhanced and reconstructed image.

[0026] like Figure 2 As shown in the figure, this application provides a method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment, including: Acquire the image sequence with polarization characteristics and the corresponding inspection pose information synchronously acquired in the tunnel inspection path; A fusion neural network with static scene branches and dynamic water mist polarization branches is constructed. The image sequence with polarization characteristics and the inspection pose information are input into the fusion neural network to decouple the scene information and output a fog-free scene image, a geometric gradient map reflecting the scene structure features, and a water mist concentration distribution map reflecting the degree of environmental degradation. A frequency domain attention generation network is constructed for extracting scene features. The geometric gradient map and the water mist concentration distribution map are input into the frequency domain attention generation network. Based on the scene geometry and water mist distribution pattern, a spatially adaptive frequency domain fusion weight map is calculated. The fog-free scene image is decomposed into a multi-scale frequency domain to obtain high-frequency components representing image details and low-frequency components representing image structure and brightness. Based on the frequency domain fusion weight map, the high-frequency components and the low-frequency components are weighted and fused at the pixel level, and inverse transformation processing is performed to generate a preliminary enhanced image. The preliminary enhanced image is then subjected to color correction and contrast equalization processing to output the final enhanced and reconstructed image.

[0027] It should be noted that the technical principle of this invention lies in constructing a phased, multimodal information fusion framework for image enhancement and reconstruction. First, an innovative fusion neural network is used to combine image sequences containing polarization characteristics with inspection pose information, achieving deep decoupling of complex scenes. This network reconstructs the fog-free tunnel's 3D geometry using the principle of motion recovery structure through a static scene branch, while simultaneously inverting the spatial distribution of water mist through a dynamic water mist polarization branch combined with a physical scattering model. This decomposes the original degraded image into three independent intermediate representations: pure scene content, geometric structure, and environmental degradation factors. Subsequently, switching to frequency domain processing, a frequency domain attention generation network intelligently analyzes the geometric structure map and water mist concentration map, generating a spatially adaptive pixel-level guidance map—the frequency domain fusion weight map—for subsequent image enhancement. This weight map precisely indicates how to specifically process detail and brightness information at each location in the image. Finally, the fog-free scene image is decomposed into high-frequency and low-frequency components, and the weight map is used for fine-grained weighted fusion and reconstruction. Subsequent color and contrast optimization further enhances the final high-quality, clear image.

[0028] In one possible implementation of this application embodiment, the method further includes: The acquired image sequence with polarization characteristics is subjected to an inter-frame consistency check to select valid image frames. Using the collected inspection pose information, the pose corresponding to the effective image frame is smoothed by filtering to generate smoothed inspection trajectory information. The image sequence composed of the effective image frames and the smoothed inspection trajectory information are used as input data for subsequent steps to improve scene decoupling and enhance the stability of reconstruction.

[0029] In some implementations, to ensure the reliability of subsequent scene decoupling and reconstruction, this invention introduces a data preprocessing stage before the core process. This preprocessing is applied to two frames with adjacent timestamps in an image sequence. and The dense optical flow field was calculated using the Farneback optical flow algorithm. Calculate the average amplitude : ; in, The total number of pixels in the image. Using pixel indices, the structural similarity index of two images in the Lab color space is calculated simultaneously. : ; in, and These represent the mean and standard deviation of the image patch, respectively. For the first The mean of the time-time graph. For the first The mean of the time-time graph. For the first The variance of the time-mapping image, For the first The variance of the time-mapping image, For the first Time and the Covariance of time-major images and The stability constant is usually taken as... , , Define the dynamic range of pixel values, such as 255 for an 8-bit image; set the motion consistency threshold. Select 5-15 pixels, structural similarity threshold Typically, it is taken as 0.80 to 0.90, when and At that time, the judgment of the first The frame is a valid image frame, thus obtaining the result from... A sequence consisting of valid frames Secondly, inspect the pose vector. The first three components are position, and the last three components are Euler angles or quaternions. For each pose component, such as A one-dimensional Savitzky-Golay filter is used for smoothing, for a length of Data points within the sliding window Least squares fitting of order-1 polynomials The value is a positive integer, either 2 or 3, corresponding to a window length of 5 or 7. Typically, the smoothed value is taken as the second or third order. The result is obtained by finding the value of the fitted polynomial at the center point of the window, and is expressed as: ; in, The weighted version of the first The output signal at time , coefficient By polynomial order and window length The unique determination can be obtained by solving the normal equations. After performing this operation independently on all pose components, a smoothed pose sequence is obtained. , For the first The original input signal at each time point; finally, the image sequence consisting of valid image frames is strictly aligned and correlated with the smoothed inspection trajectory information according to the timestamp. This correlation ensures that each valid image frame... Each has a unique corresponding, noise-reduced pose. It effectively suppresses reconstruction instability and artifacts caused by data acquisition device jitter, momentary occlusion, or sensor noise.

[0030] For example, suppose that during a tunnel inspection, the inspection robot simultaneously acquires two grayscale images with polarization characteristics that have adjacent timestamps. and All of them are set to resolution First, the Farneback optical flow algorithm is used to calculate... Compared to Dense optical flow field The sum of the displacement magnitudes of all pixels was calculated to be: Then the average motion offset is Pixels. Simultaneously, the image is converted to the Lab color space to calculate structural similarity; let the mean value obtained at this time be... , Standard deviation , covariance Set dynamic range Then the stability constant , Substituting into the formula, we obtain... Set a motion consistency threshold. Pixels, structural similarity threshold ,because and Then determine the first The frame is a valid image frame. For inspection pose information, it is assumed that the original image frame acquired is... Axis position sequence in Within the neighborhood of time Meters, using a window length of 5, that is order The filter is smoothed, and the corresponding filter coefficients are... The smoothed position value Meters. This will make the effective frame... With smooth pose Strict alignment and association.

[0031] In one possible implementation, combining Figure 2 The construction of the fusion neural network with static scene branches and dynamic water mist polarization branches includes: Using a tunnel image degradation simulator incorporating a differentiable physical scattering model, a simulated polarization image sequence and corresponding simulated water mist concentration map are generated based on clear tunnel images; The simulated polarization image sequence is input into the fusion neural network to be trained. The backpropagation algorithm is used to minimize the image reconstruction loss between the fog-free scene image and the clear tunnel image, as well as the concentration estimation loss between the output water mist concentration distribution map and the simulated water mist concentration map, to complete the network parameter update.

[0032] In some implementations, a tunnel image degradation simulator incorporating a differentiable physical scattering model is used for a given sharp tunnel scene image. ,in Represents pixel coordinates and the water mist density field distributed in three-dimensional space. , Representing three-dimensional spatial coordinates, the simulation of the degradation process follows an improved atmospheric scattering model and polarization transport equation, simulating the generation at a polarization angle. Degraded images acquired below It can be represented as: ; in, This is a transmittance diagram, which describes the attenuation of light along its propagation path. It is the total scattering coefficient related to ambient humidity, with units of . , It is a pixel The distance from the corresponding scene point to the camera, and the integral path. Along the line of sight, This is the global atmospheric light value; inside tunnels, the color temperature of artificial light sources is typically used. It is a term characterizing the polarization modulation effect, and it is related to the polarization angle. Related to the polarization attenuation characteristics of water mist, it can be modeled as follows: ; in, It is a polarization attenuation coefficient related to the size distribution of water mist particles, with a value ranging from 0 to 1. It refers to the polarization angle distribution of the reflected light from the scene; by setting different... , usually ( This allows for the batch generation of simulated polarization image sequences from multiple angles. and the simulated water mist concentration graph as the true value Secondly, the simulated polarization image sequence is input into the fusion neural network to be trained, and the backpropagation algorithm is used to minimize the fog-free scene image output by the network. With clear tunnel images Image reconstruction loss between And the water mist concentration distribution map that minimizes the network output. Comparison with simulated water mist concentration diagram Concentration estimation loss between This completes the network parameter update; image reconstruction loss. A hybrid loss function combining L1 norm and structural similarity is employed: ; in, and These are the hyperparameters for balancing the weights, set to 1.0 and 0.2 respectively; This is a structural similarity index that measures the similarity between the predicted image and the real image in terms of brightness, contrast, and structure. Its value ranges from [value range missing]. The closer the value is to 1, the more similar the two are; Employ smoothed L1 loss for greater robustness to outliers: ; in, This is a threshold parameter, usually set to 1.0; Represents the total number of image pixels involved in the calculation; total loss , The weighting factor for concentration loss is typically 0.5; the total loss is calculated using standard stochastic gradient descent or the Adam optimizer. Relative to all network weight parameters gradient According to the learning rate ,like Update parameters: The iteration continues until the loss converges.

[0033] For example, training data is generated using a tunnel image degradation simulator, and a clear tunnel background image is selected. Set its global atmospheric light value The value is 0.85, corresponding to the normalized value of the sodium lamp color temperature inside the tunnel. This assumes a non-uniform water mist density field generated in the simulation. At a certain pixel The integral result along the line of sight, i.e., the simulated water mist concentration at that point. The value is 0.45. Let the overall environmental scattering coefficient be... polarization attenuation coefficient Polarization angle distribution When simulating the acquisition of polarization angle When calculating the transmittance of a degraded image, Calculate the polarization modulation term If the pixel resolution value The generated simulated polarization image pixel values During training, the generated sequences are input into a fusion neural network, assuming the network output predicts the pixel values ​​of the haze-free image. Predict water mist concentration Set the weights for the loss function. Calculate the image reconstruction loss. Suppose at this time The value is 0.98, and the concentration is used to estimate the loss. Finally, calculate the total loss. The Adam optimizer adjusts the network weights based on this loss. Perform backpropagation update, learning rate Set as This allows the network to output in subsequent iterations. and Approaching the true value and .

[0034] In one possible implementation, combining Figure 2 The process of inputting the image sequence with polarization characteristics and the inspection pose information into a fusion neural network for scene information decoupling includes: Stokes vector calculation is performed on the image sequence with polarization characteristics collected on site to obtain a polarization degree map reflecting the intensity of water mist scattering and a polarization angle map reflecting the vibration direction of polarized light. The polarization degree map, the polarization angle map and the inspection pose information are used as the joint input data of the trained fusion neural network. Through the static scene branch, the tunnel scene is implicitly modeled in three dimensions based on the inspection pose information, a fog-free scene image is rendered and output, and the scene geometric gradient map is derived by taking the derivative of the implicit function. By using the dynamic water mist polarization branch, a physical modulation model of water mist on polarized light is learned, and physical consistency constraints are imposed by combining the polarization degree diagram and the polarization angle diagram to invert and output the water mist concentration distribution map.

[0035] In some implementations, Stokes vector calculations are performed on image sequences with polarization characteristics to obtain a polarization degree map reflecting the intensity of water mist scattering and a polarization angle map reflecting the vibration direction of polarized light. For each pixel, it is assumed that intensity images with four different polarization angles were acquired. , usually Then the Stokes vector Calculated using the following system of linear equations: ; For linearly polarized light It is usually assumed to be 0; degree of polarization and polarization angle They can be calculated separately as follows: ; ; in, The arctangent function outputs the angle at... Between radians, generate a polarization map of the same size as the original image. and polarization angle diagram ; Polarization degree diagram, polarization angle diagram, and inspection pose information The joint input data is used for the trained fusion neural network; through a static scene branch, a 3D implicit model of the tunnel scene is performed based on the inspection pose information, and a fog-free scene image is rendered and output. An architecture similar to a neural network radiation field is adopted, using a multilayer perceptron. The three-dimensional coordinates after position encoding and view direction Mapped to volume density and color value , , It is a high-frequency position coding function. , This is the number of coding frequencies, usually taken as 10. These are normalized coordinate components; given a series of known camera poses That is, the inspection pose is obtained by sampling along the ray of each pixel and synthesizing the color at that viewpoint using the volume rendering integral formula: ; in, It is the cumulative transmittance. It is the ray parameter equation. It is the optical center of the camera, and the integral is near the boundary. and distant boundary This process is performed between steps, using discretized sampling and differentiable rendering. This branch outputs a fog-free scene image. Meanwhile, through implicit functions Regarding spatial coordinates Differentiate and derive the geometric gradient map of the scene. , Represents the spatial gradient operator. It is a pixel The corresponding key surface points are typically determined by the density field peaks, and the gradient map reflects the geometric changes of the scene surface. Through a dynamic water mist polarization branch, a physical modulation model of the water mist on polarized light is learned. Physical consistency constraints are applied by combining the degree of polarization map and the polarization angle map, and the water mist concentration distribution map is inverted and output. Geometric information from the static branch and the input are also considered. and Modeling a 3D water mist density field and polarization modulation parameter field Then, through a differentiable volume rendering process, the change in polarization state of light after passing through the water mist field is simulated, and the predicted polarization degree map is rendered. and polarization angle diagram And by calculating the physical consistency loss: ; in, The branch parameters are optimized using angle loss weights. , Taking 0.5, the optimized three-dimensional water mist density field is finally... Integrating along the viewing direction of each pixel yields a two-dimensional water mist concentration distribution map. .

[0036] For example, firstly, four frames with different polarization angles are obtained ( ) performs pixel-level analysis on the intensity image, assuming a certain pixel point The corresponding strength values ​​are respectively The Stokes vector components at that point are calculated using a system of linear equations to obtain the total intensity. , differential component as well as Then, the polarization degree of that pixel is calculated. and polarization angle The generated polarization map Polarization angle diagram and the corresponding smooth pose After inputting into the fusion neural network, the static scene branch utilizes the positional encoding function. Process the coordinates of this point and set the encoding frequency. Through multilayer perceptron The volume density of this spatial point is obtained by mapping. and color value Volume rendering is performed along the line of sight, assuming cumulative translucency. Step length The rendered predicted haze-free image pixel values The geometric gradient is derived by differentiating the implicit function. Simultaneously, a dynamic water mist polarization branch learning physical modulation model is used, assuming the estimated water mist density... Under the constraint of physical consistency, the predicted polarization characteristics and measured values ​​are calculated. If the predicted value of the L1 loss is 0.20, then the polarization physical loss component is... Update parameters by minimizing this physical consistency loss. Finally, the water mist concentration at that point is obtained by integrating along the line of sight. .

[0037] In one possible implementation, combining Figure 2 The construction of the frequency domain attention generation network for extracting scene features includes: An operator for edge enhancement processing of the scene's geometric gradient map and a mapping unit for nonlinear mapping processing of the water mist concentration distribution map are constructed. A multi-channel stitching layer is established at the output of the operator and the mapping unit, and several convolutional layers are connected after the multi-channel stitching layer to fuse the detail preservation weight map generated by the operator with the fog suppression weight map generated by the mapping unit to extract a composite feature map that reflects the scene structure and the distribution law of water mist. An attention mechanism layer containing two parallel convolutional branches is set after the convolutional layer. The two parallel convolutional branches are respectively configured to calculate the high-frequency attention score and low-frequency attention score of each pixel position based on the composite feature map, forming a spatially adaptive frequency domain fusion weight map. Using the simulated water mist concentration map generated by the tunnel image degradation simulator as training data, the parameters of the frequency domain attention generation network are updated by calculating the correlation loss between the weight distribution output by the frequency domain attention generation network and the simulated water mist concentration map.

[0038] In some implementations, an operator is constructed for edge enhancement processing of the scene's geometric gradient map, and a mapping unit is constructed for nonlinear mapping processing of the water mist concentration distribution map. The edge enhancement operator employs a learnable convolutional kernel. The size is , for the input geometric gradient map Perform convolution operations to enhance edge response and suppress noise, generating a detail-preserving weight map. : ; in, This is a two-dimensional convolution operation. It is the Sigmoid activation function, which normalizes the output to the (0,1) interval. The expression is: Simultaneously, the nonlinear mapping unit outputs the water mist concentration distribution map from the input. A nonlinear transformation is performed to generate a fog suppression weight map. : ; in, and Indicates a fully connected layer or Convolutional layer It is the modified linear unit activation function, and its expression is: A multi-channel splicing layer is established at the output of the operator and mapping unit to... and The feature maps are stitched together along the channel dimension to form a dual-channel feature map. ,in This refers to connecting multiple tensors together along a specific dimension to form a larger tensor, and then connecting several convolutional layers, such as three layers, after a multi-channel splicing layer. Convolutional layers, followed by ReLU activation, with 32, 64, and 128 channels respectively, are used to fuse the detail preservation weight map generated by the operator with the fog suppression weight map generated by the mapping unit, extracting a composite feature map reflecting the scene structure and water mist distribution patterns. ,Right now An attention mechanism layer containing two parallel convolutional branches is set after the convolutional layer. The two parallel convolutional branches are configured to be based on composite feature maps. Calculate the high-frequency attention score for each pixel location. and low-frequency attention scores Each branch is composed of It consists of convolutional layers and a sigmoid activation function, i.e. and ,in, and It is independent The convolutional layer maps the 128-channel composite features into single-channel attention maps. These two score maps together form a spatially adaptive frequency domain fusion weight map. The size is Pixel-level fusion weights corresponding to high-frequency and low-frequency components, respectively; simulated water mist concentration map generated using a tunnel image degradation simulator. As training data, the parameters of the frequency domain attention generation network are updated by calculating the correlation loss between the weight distribution of the network's output and the simulated water mist concentration map. The correlation loss is defined as the high-frequency attention score. Compared with simulated concentration map The negative correlation, and low-frequency attention scores Compared with simulated concentration map The positive correlation, the specific loss function is: ; in, Calculate the Pearson correlation coefficient between the two graphs. It is The result normalized to the [0,1] interval and These are the weighting coefficients for balancing the two terms, both set to 1.0. The goal is to minimize this correlation loss. The Adam optimizer is used for backpropagation, and the learning rate is set to... Update all parameters in the frequency domain attention generation network. Figure 3 The image illustrates the weight distribution characteristics of the frequency domain attention generation network output. The high-contrast regions in the figure reflect the sensitive response of the edge enhancement operator to the geometric gradient, while the gray-level variations demonstrate the weight adjustment logic of the nonlinear mapping unit for different water mist concentration regions, proving the spatial adaptability of the network in the complex environment of tunnels.

[0039] For example, obtain the decoupled scene geometry gradient map. Water mist concentration distribution map Set a specific pixel. Geometric gradient at water mist concentration The edge reinforcement operator adopts... convolution kernel The process is performed, assuming the response value after convolution is 1.20, and then the Sigmoid activation function is used to calculate the guard-of-mind weights. Meanwhile, the nonlinear mapping unit processes the concentration values ​​through a fully connected layer. Assuming , and The processed output value is -0.85. The fog suppression weight is then calculated. The obtained weights are concatenated and fused through a convolutional layer to generate a 128-channel composite feature map. Two parallel convolutional branches calculate the high-frequency and low-frequency attention scores respectively. Assuming the convolutional output values ​​at corresponding pixel positions are 0.65 and -0.40 respectively, the high-frequency attention score... Low-frequency attention score During the training phase, simulated ground truth is used. Calculate the Pearson correlation coefficient, assuming Set weights Then calculate the correlation loss. The Adam optimizer uses this loss to... The learning rate updates the network parameters, enabling the network to automatically generate a spatially adaptive frequency domain fusion weight map based on the geometric gradient and concentration distribution. .

[0040] In one possible implementation, combining Figure 2 The differentiable physical scattering model includes: The differentiable physical scattering model includes a scattering coefficient related to ambient humidity and a polarization attenuation coefficient related to the size of water mist particles. During the execution of the backpropagation algorithm, the values ​​of the physical parameters are dynamically adjusted by calculating the gradient of the total loss relative to the scattering coefficient and the polarization attenuation coefficient. By updating the physical parameters, the feature distribution of the simulated polarization image sequence is adjusted in a direction that makes the simulation degradation closer to the distribution of real acquired data, thereby enhancing the generalization ability of the fused neural network and frequency domain attention generation network.

[0041] In some implementations, the differentiable physical scattering model includes scattering coefficients related to ambient humidity. And the polarization attenuation coefficient related to the size of water mist particles. , and Defined as a trainable tensor, it participates in gradient backpropagation throughout the training process as part of the tunnel image degradation simulator. During the backpropagation algorithm execution, it is quantized by the image reconstruction loss. and concentration estimation loss Composition, by calculating total loss If training is involved, the total loss may also include correlation loss. Relative to scattering coefficient and polarization attenuation coefficient gradient and The values ​​of these two physical parameters are dynamically adjusted. The formula for updating the parameters is: ; ; ; ; in, and It is an estimate of the first and second moments of the gradient. The corrected first-moment estimate represents an unbiased estimate of the exponentially weighted moving average of the parameter gradient. The corrected second-moment estimate represents an unbiased estimate of the exponentially weighted moving average of the squared gradients of the parameters. and This is the exponential decay rate estimated by moments, typically taken as 0.9 or 0.999. It is the learning rate, which is generally taken as... , It is the numerical stability constant, and is generally taken as... superscript Indicates the number of iterations. For The updates use the exact same formula. Update physical parameters. and To simulate polarization image sequences The feature distribution is adjusted in a direction that makes the simulation degradation more closely resemble the distribution of the real collected data. Gradient Instructions on how to adjust the scattering intensity to make the simulated degradation's overall contrast and brightness closer to a real foggy image, while the gradient... It indicates how to adjust the polarization modulation characteristics to produce a simulated polarization degree map. and polarization angle diagram The statistical distribution is closer to the polarization diagram calculated from real data. Figure 4 The scattering coefficient was recorded. With polarization attenuation coefficient The numerical evolution curves during iterative training. The convergence trend of the curves indicates that the differentiable update mechanism based on the Adam optimizer enables the physical parameters of the degradation simulator to automatically match the distribution characteristics of the real-world collected data, thereby effectively improving the generalization performance of the overall algorithm.

[0042] For example, suppose that during the training initialization phase, the scattering coefficients in the tunnel image degradation simulator... Initially set to polarization attenuation coefficient The initial value is set to 0.50. In one training iteration, the total loss is calculated. The partial derivatives with respect to the physical parameters, assuming the resulting gradient values ​​are as follows: and The Adam optimizer is used for updates, with an exponential decay rate set. Learning rate Numerical stability constant .by Taking the update as an example, suppose the first moment of the previous time step... and second moment If both are 0, then the first moment at the current time is... Second moment After deviation correction, Substituting into the formula, the scattering coefficient at the next time step... Similarly, if The calculated update increment is 0.002, therefore the updated... Through this dynamic adjustment mechanism, the gradient Guided scattering intensity adjustment makes the contrast of the simulated image closer to that of a real foggy image, gradient This guides the polarization distribution to closely approximate the actual polarization characteristics, thereby enabling the simulator's feature distribution to continuously approach the actual tunnel acquisition data.

[0043] In one possible implementation, combining Figure 2 The process involves pixel-level weighted fusion of the high-frequency and low-frequency components, followed by inverse transform processing to generate a preliminary enhanced image, including: From the spatially adaptive frequency domain fusion weight graph of the frequency domain attention generation network output, the high-frequency channel weights corresponding to the high-frequency components and the low-frequency channel weights corresponding to the low-frequency components are extracted respectively. The high-frequency channel weights are multiplied pixel by pixel with the high-frequency components to obtain a weighted high-frequency sub-band image, which preserves details in the edge region and suppresses noise in the flat region. The low-frequency channel weights are multiplied pixel by pixel with the low-frequency components to obtain a weighted low-frequency sub-band image, which can restore illumination in dense fog areas and avoid overexposure in well-lit areas. The weighted high-frequency sub-band image and the weighted low-frequency sub-band image are added pixel by pixel to obtain the fused frequency domain sub-band image; A multi-scale inverse frequency domain transform is performed on the fused frequency domain sub-band image to recombine the frequency domain features and transform them back into the spatial domain, generating a preliminary enhanced image.

[0044] In some implementations, a spatially adaptive frequency-domain fusion weight graph is generated from the output of the frequency-domain attention network. In the middle, the size is ,in and The high-frequency channel weights corresponding to the high-frequency components are extracted from the image height and width, respectively. and the low-frequency channel weights corresponding to the low-frequency components The values ​​of both images are in the range of (0,1), representing the enhancement intensity of high-frequency details and low-frequency structures at each pixel location, respectively; High-frequency components representing image details obtained through multi-scale frequency domain decomposition Perform pixel-by-pixel multiplication to obtain the weighted high-frequency subband image: ; in, Representing pixel coordinates, the amplitude of high-frequency components is adaptively adjusted according to the weight map to preserve or even enhance detailed textures in edge regions, while effectively suppressing noise amplification in flat or noisy areas; Low-frequency components representing image structure and brightness Perform pixel-by-pixel multiplication to obtain the weighted low-frequency sub-band image: ; By adjusting the intensity of low-frequency components based on the weighted graph, illuminance can be increased in dense fog areas to restore the scene's intrinsic brightness. This typically corresponds to higher water fog concentrations. Larger values, for example In well-lit or recovered areas, avoid overexposure due to excessive enhancement; then, the weighted high-frequency subband image... Weighted low-frequency sub-band image By adding pixels one by one, the fused frequency domain sub-band image is obtained. The image contains complete image information in the frequency domain after spatial adaptive modulation; the fused frequency domain sub-band image Perform a multi-scale inverse frequency domain transform to recombine and transform the frequency domain features back into the spatial domain, generating a preliminary enhanced image. Multiscale frequency domain inverse transform is the inverse operation of the previous decomposition process, such as using Laplace pyramids or wavelet transforms. For a layered Laplacian pyramid, the inverse transform process starts from the low-frequency sub-band at the top layer, and upsamples and adds it layer by layer with the corresponding high-frequency sub-band until the original resolution is restored. Specifically, let the reconstructed image of the l-th layer be... ,initialization The lowest frequency at the top layer, for ,implement ,in This indicates an upsampling operation, such as bilinear interpolation. It is the high-frequency component after weighting at layer l, and finally This is the initial enhanced image generated.

[0045] For example, the size of the output of the frequency domain attention generation network is Spatial Adaptive Frequency Domain Fusion Weighting Graph In the process, the high-frequency channel weight map is extracted. Low-frequency channel weighting diagram Suppose that for a pixel in the edge region with coordinates (960, 540), its high-frequency weights... Low-frequency weights For pixels in a flat, dense fog region with coordinates (100, 100), the high-frequency weights... Low-frequency weights During the fusion process, let the original high-frequency component value at (960, 540) be... Low-frequency component value The weighted high-frequency subband value is The weighted low-frequency subband value is The frequency domain subband value after fusion at this point is This achieves enhanced preservation of edge details. For the point (100, 100), let its original value be... Then after weighting The fusion value is This effectively suppressed noise and improved illumination in foggy areas. Finally, a multi-scale inverse frequency domain transform was performed, assuming the following was adopted. The pyramid of Laplace, from the top layer Initially, upsampling is performed layer by layer, and the high-frequency components of the corresponding layer are added back, as in the following process. Until the original resolution is restored. This yields the final preliminary enhanced image. .

[0046] In one possible implementation, combining Figure 2 The preliminary enhanced image undergoes color correction and contrast equalization processing to output the final enhanced and reconstructed image, including: The preliminary enhanced image is converted from the RGB color space to the HSV color space, and the luminance component map is extracted from it. The brightness component map is subjected to contrast-limited adaptive histogram equalization to enhance the local contrast of the image while suppressing noise amplification, thereby generating equalized brightness components. The equalized luminance component is merged with the original hue and saturation components in the HSV color space, and the merged data is converted back to the RGB color space to output the final enhanced reconstructed image.

[0047] In some implementations, the image will be initially enhanced. Convert from RGB color space to HSV color space, where HSV represents hue, saturation, and lightness, respectively. Pixel values ​​in the RGB color space are typically normalized to the range [0,1]. The conversion formula follows the standard color space conversion model, applying the RGB values ​​of each pixel. Its corresponding The component, namely the luminance component, is calculated as follows: Saturation The calculation formula is: And the color tone The calculation involves piecewise functions, let's say... ,but ,final It is normalized to the [0,1] interval. The luminance component image is then extracted using this transformation. For the luminance component map Perform contrast-limited adaptive histogram equalization to enhance local image contrast while suppressing noise amplification, generating equalized luminance components. The core idea is to divide the image into several non-overlapping local regions, such as tiles, with a size of [missing information]. For each pixel, its luminance histogram is calculated independently. Then, each histogram is cropped, and those exceeding the preset cropping limit are removed. The cropped histogram is uniformly redistributed across all histogram intervals to prevent noise amplification caused by excessive local contrast enhancement. The cumulative distribution function of the cropped histogram is calculated and normalized to obtain the mapping function for that tile. Finally, bilinear interpolation is used to smoothly blend the mapping functions of adjacent tiles and apply them to each pixel, resulting in a luminance component with globally balanced contrast and enhanced local details. The essence is to find the mapping. This results in an approximately uniform distribution of output brightness within a local area. The brightness component is then equalized. Compared to the original hue components in the HSV color space and saturation components Merge the images to create a new HSV image. The merged data Convert back to RGB color space to output the final enhanced reconstructed image. Inverse transformation: Multiply by the required amount first. Convert back to angle , and Calculate intermediate variables, let .according to The interval is determined Value, for example when hour, ;when hour, Finally obtained ,Will The value is limited to the range [0,1].

[0048] For example, suppose the generated preliminary enhanced image The normalized RGB value of a certain pixel in the image First, convert it to HSV color space and calculate the luminance component. ; Calculate the saturation component, because ,but ; Calculate hue components Because of the maximum value Difference but After normalization Extracting the luminance component image. Then perform CLAHE processing, assuming the pixel is located in Local region histogram After cropping and redistribution, the corresponding cumulative distribution function mapping value is 0.72. The equalized luminance component is then calculated using bilinear interpolation. Then Merge and convert back to RGB space, calculate intermediate variables ,because Corresponding angle In Interval, calculation Calculate the brightness compensation amount The final RGB reconstructed value for that point is... That is, output the final enhanced and reconstructed image. The value at that point .

[0049] In one possible implementation, combining Figure 2 By using the dynamic water mist polarization branch, a physical modulation model of water mist on polarized light is learned, and physical consistency constraints are applied by combining the polarization degree map and polarization angle map. The resulting water mist concentration distribution map is inverted and output as follows: Based on the inspection pose information and the scene surface distance information generated by the static scene branch, the three-dimensional space within the current camera's view frustum is sampled to estimate the water mist density at each sampling point and the polarization modulation parameters characterizing the effect of water mist particles on the polarization state of light. Using the estimated water mist density and polarization modulation parameters at each sampling point, combined with the fog-free scene image, the volume rendering equation is used to simulate the change in polarization characteristics of light passing through the water mist, and a predicted polarization degree map and a predicted polarization angle map corresponding to the current viewpoint are generated. The difference between the predicted polarization degree map and the actual calculated polarization degree map, and the difference between the predicted polarization angle map and the actual calculated polarization angle map are calculated to obtain the polarization physical loss; The weights of the dynamic water mist polarization branch are updated using the polarization physical loss, and the optimized three-dimensional water mist density field is integrated along the line of sight to generate a two-dimensional water mist concentration distribution map.

[0050] In some implementations, image inpainting is achieved through data preprocessing, scene decoupling, frequency domain weight calculation, and multi-scale fusion. In the preprocessing stage, adjacent images are processed... and The average motion offset was calculated using the Farneback optical flow algorithm. And combined with the structural similarity index in Lab space Filter valid frames and simultaneously process the pose vectors. One-dimensional Savitzky-Golay filtering is performed to smooth the data and eliminate physical jitter interference during the inspection process. The static branch utilizes the NeRF architecture to convert the three-dimensional coordinates... and view direction Mapped to volume density and color A fog-free image is synthesized using a volume rendering equation, and a geometric gradient map is derived. The dynamic branch estimates the water mist density using a multilayer perceptron. and polarization parameters Rendering generates predicted Stokes vectors This leads to the inversion of water mist concentration. Frequency domain attention generation networks utilize convolutional kernels. Calculate detail protection weights Combined with dense fog suppression weights Extracting composite features A loss function is constructed using the Pearson correlation coefficient. Parameters are updated. During the fusion phase, the high-frequency subbands are adjusted according to the weight map. and low-frequency subband Perform weighted summation using the inverse transformation formula. Image restoration. Brightness equalization is performed in HSV space and then reversed back to RGB space to improve visual quality.

[0051] For example, the camera optical center position is determined based on the inspection pose information. and line of sight Combined with the scene surface distance generated by the static scene branch meters, in near-plane Meters and distant plane Select between meters along the line of sight Each sampling point. For a given sampling point Assuming dynamic water mist polarization branch is used to estimate its water mist density And polarization attenuation coefficient The volume rendering equation is used to simulate the change in polarization characteristics. If the brightness of the corresponding fog-free scene for that pixel... Step length Cumulative transmittance Incident polarization angle And the inherent polarization angle of the water mist The contribution of this sampling point to the prediction of the Stokes vector is calculated as follows: Components ; Components ; Components Assume the predicted polarization feature obtained after accumulating all sampling points is: and The true value calculated from the measured image is... and Set weights The polarization physical loss was calculated. This loss is used to update the branch weights using the Adam optimizer. Then, the optimized three-dimensional density field is discretized and integrated along the line of sight, as shown in the calculation. If the average density at each point is 0.045, then the value of the final generated two-dimensional water mist concentration distribution map at that pixel is... .

[0052] It should be noted that all equivalent changes and modifications made in accordance with the teachings of this invention are still within the scope of this invention. Those skilled in the art will readily conceive of other embodiments of this invention upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this invention that follow the general principles of this invention and include common knowledge or conventional techniques in the art not described herein.

Claims

1. A method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment, characterized in that, The method includes: Acquire the image sequence with polarization characteristics and the corresponding inspection pose information synchronously acquired in the tunnel inspection path; A fusion neural network with static scene branches and dynamic water mist polarization branches is constructed. The image sequence with polarization characteristics and the inspection pose information are input into the fusion neural network to decouple the scene information and output a fog-free scene image, a geometric gradient map reflecting the scene structure features, and a water mist concentration distribution map reflecting the degree of environmental degradation. A frequency domain attention generation network is constructed for extracting scene features. The geometric gradient map and the water mist concentration distribution map are input into the frequency domain attention generation network. Based on the scene geometry and water mist distribution pattern, a spatially adaptive frequency domain fusion weight map is calculated. The fog-free scene image is decomposed into a multi-scale frequency domain to obtain high-frequency components representing image details and low-frequency components representing image structure and brightness. Based on the frequency domain fusion weight map, the high-frequency components and the low-frequency components are weighted and fused at the pixel level, and inverse transformation processing is performed to generate a preliminary enhanced image. The preliminary enhanced image is then subjected to color correction and contrast equalization to output the final enhanced and reconstructed image.

2. The method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment according to claim 1, characterized in that, The method further includes: The acquired image sequence with polarization characteristics is subjected to an inter-frame consistency check to select valid image frames. Using the collected inspection pose information, the pose corresponding to the effective image frame is smoothed by filtering to generate smoothed inspection trajectory information. The image sequence composed of the effective image frames and the smoothed inspection trajectory information are used as input data for subsequent steps to improve scene decoupling and enhance the stability of reconstruction.

3. The method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment according to claim 2, characterized in that, The construction of the fusion neural network with static scene branches and dynamic water mist polarization branches includes: Using a tunnel image degradation simulator incorporating a differentiable physical scattering model, a simulated polarization image sequence and corresponding simulated water mist concentration map are generated based on clear tunnel images; The simulated polarization image sequence is input into the fusion neural network to be trained. The backpropagation algorithm is used to minimize the image reconstruction loss between the fog-free scene image and the clear tunnel image, as well as the concentration estimation loss between the output water mist concentration distribution map and the simulated water mist concentration map, to complete the network parameter update.

4. The method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment according to claim 3, characterized in that, The process of inputting the image sequence with polarization characteristics and the inspection pose information into a fusion neural network for scene information decoupling includes: Stokes vector calculation is performed on the acquired image sequence with polarization characteristics to obtain a polarization degree map reflecting the intensity of water mist scattering and a polarization angle map reflecting the vibration direction of polarized light. The polarization degree map, the polarization angle map and the inspection pose information are used as the joint input data of the trained fusion neural network. Through the static scene branch, the tunnel scene is implicitly modeled in three dimensions based on the inspection pose information, a fog-free scene image is rendered and output, and the scene geometric gradient map is derived by taking the derivative of the implicit function. By using the dynamic water mist polarization branch, a physical modulation model of water mist on polarized light is learned, and physical consistency constraints are applied by combining the polarization degree diagram and polarization angle diagram to invert and output the water mist concentration distribution map.

5. The method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment according to claim 3, characterized in that, The construction of the frequency domain attention generation network for extracting scene features includes: An operator for edge enhancement processing of the scene's geometric gradient map and a mapping unit for nonlinear mapping processing of the water mist concentration distribution map are constructed. A multi-channel stitching layer is established at the output of the operator and the mapping unit, and several convolutional layers are connected after the multi-channel stitching layer to fuse the detail preservation weight map generated by the operator with the fog suppression weight map generated by the mapping unit to extract a composite feature map that reflects the scene structure and the distribution law of water mist. An attention mechanism layer containing two parallel convolutional branches is set after the convolutional layer. The two parallel convolutional branches are respectively configured to calculate the high-frequency attention score and low-frequency attention score of each pixel position based on the composite feature map, forming a spatially adaptive frequency domain fusion weight map. Using the simulated water mist concentration map generated by the tunnel image degradation simulator as training data, the parameters of the frequency domain attention generation network are updated by calculating the correlation loss between the weight distribution output by the frequency domain attention generation network and the simulated water mist concentration map.

6. A method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment according to claim 3 or 5, characterized in that, The differentiable physical scattering model includes: The differentiable physical scattering model includes a scattering coefficient related to ambient humidity and a polarization attenuation coefficient related to the size of water mist particles. During the execution of the backpropagation algorithm, the values ​​of the physical parameters are dynamically adjusted by calculating the gradient of the total loss relative to the scattering coefficient and the polarization attenuation coefficient. By updating the physical parameters, the feature distribution of the simulated polarization image sequence is adjusted in a direction that makes the simulation degradation closer to the distribution of real acquired data, thereby enhancing the generalization ability of the fused neural network and frequency domain attention generation network.

7. The method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment according to claim 2, characterized in that, The process of performing pixel-level weighted fusion of the high-frequency components and the low-frequency components, followed by inverse transform processing, to generate a preliminary enhanced image includes: From the spatially adaptive frequency domain fusion weight map output by the frequency domain attention generation network, high-frequency channel weights corresponding to high-frequency components and low-frequency channel weights corresponding to low-frequency components are extracted respectively. The high-frequency channel weights are multiplied pixel by pixel with the high-frequency components to obtain a weighted high-frequency sub-band image, which preserves details in the edge region and suppresses noise in the flat region. The low-frequency channel weights are multiplied pixel by pixel with the low-frequency components to obtain a weighted low-frequency sub-band image, which can restore illumination in dense fog areas and avoid overexposure in well-lit areas. The weighted high-frequency sub-band image and the weighted low-frequency sub-band image are added pixel by pixel to obtain the fused frequency domain sub-band image; A multi-scale inverse frequency domain transform is performed on the fused frequency domain sub-band image to recombine the frequency domain features and transform them back into the spatial domain, generating a preliminary enhanced image.

8. The method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment according to claim 2, characterized in that, The preliminary enhanced image is subjected to color correction and contrast equalization processing to output the final enhanced and reconstructed image, including: The preliminary enhanced image is converted from the RGB color space to the HSV color space, and the luminance component map is extracted from it. The brightness component map is subjected to contrast-limited adaptive histogram equalization to enhance the local contrast of the image while suppressing noise amplification, thereby generating equalized brightness components. The equalized luminance component is merged with the original hue and saturation components in the HSV color space, and the merged data is converted back to the RGB color space to output the final enhanced reconstructed image.

9. The method for descattering enhancement and reconstruction of inspection images in a tunnel water mist environment according to claim 4, characterized in that, By employing the dynamic water mist polarization branch, a physical modulation model of the water mist on polarized light is learned. Combined with the polarization degree map and polarization angle map, physical consistency constraints are applied, and the resulting water mist concentration distribution map is inverted and output, including: Based on the inspection pose information and the scene surface distance information generated by the static scene branch, the three-dimensional space within the current camera's view frustum is sampled to estimate the water mist density at each sampling point and the polarization modulation parameters characterizing the effect of water mist particles on the polarization state of light. Using the estimated water mist density and polarization modulation parameters at each sampling point, combined with the fog-free scene image, the volume rendering equation is used to simulate the change in polarization characteristics of light passing through the water mist, and a predicted polarization degree map and a predicted polarization angle map corresponding to the current viewpoint are generated. The difference between the predicted polarization degree map and the actual calculated polarization degree map, and the difference between the predicted polarization angle map and the actual calculated polarization angle map are calculated to obtain the polarization physical loss; The weights of the dynamic water mist polarization branch are updated using the polarization physical loss, and the optimized three-dimensional water mist density field is integrated along the line of sight to generate a two-dimensional water mist concentration distribution map.

10. A system for descattering, enhancing, and reconstructing inspection images in a tunnel water mist environment, characterized in that, The system includes: Synchronous acquisition module: used to acquire image sequences with polarization characteristics and corresponding inspection pose information synchronously acquired in the tunnel inspection path; Scene information decoupling module: used to construct a fusion neural network with static scene branch and dynamic water mist polarization branch, input the image sequence with polarization characteristics and the inspection pose information into the fusion neural network to decouple the scene information, and output fog-free scene image, geometric gradient map reflecting scene structural features and water mist concentration distribution map reflecting the degree of environmental degradation; Frequency domain weight calculation module: used to construct a frequency domain attention generation network for extracting scene features. The geometric gradient map and the water mist concentration distribution map are input into the frequency domain attention generation network. Based on the scene geometry and water mist distribution pattern, a spatially adaptive frequency domain fusion weight map is calculated. Multi-scale frequency domain decomposition module: used to perform multi-scale frequency domain decomposition on the fog-free scene image to obtain high-frequency components representing image details and low-frequency components representing image structure and brightness; Pixel-level weighted fusion module: used to perform pixel-level weighted fusion of the high-frequency components and the low-frequency components according to the frequency domain fusion weight map, and perform inverse transformation processing to generate a preliminary enhanced image; Visual quality optimization module: used to perform color correction and contrast equalization on the preliminary enhanced image, and output the final enhanced and reconstructed image.