AOD-Net low-exposure monitoring image defogging method based on self-attention mechanism

By using hollow convolution, Swin Transformer and RRDNet modules in the AOD-Net model, feature capture and lighting decomposition are enhanced, and the problem of image quality degradation in dynamic haze scenarios is solved, achieving efficient fog removal and brightness improvement.

CN120339110APending Publication Date: 2025-07-18CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510503700.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has poor fog removal effect when dealing with dynamic haze scenes. The AOD-Net model has limitations in detail recovery and brightness retention, especially in low exposure conditions.

Method used

The standard convolution is replaced by hollow convolution, the Swin Transformer module and RRDNet module are added, and the feature capture capability is enhanced through the self-attention mechanism, combined with the Retinex model to decompose lighting and noise, and iteratively optimize image recovery.

Benefits of technology

It improves the clarity and brightness of the image, reduces the blur caused by fog occlusion, and the restored image is more natural, conforms to the visual experience of the human eye, and improves the image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339110A_ABST
    Figure CN120339110A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to an AOD-Net low-exposure monitoring image defogging method based on a self-attention mechanism, which comprises the following steps: S1, designing dilated convolution to replace standard convolution in AOD-Net, totally five convolution layers exist in a network, utilizing the characteristics of the dilated convolution to respectively adjust the voidage of the five dilated convolution layers, and calculating the voidage of the dilated convolution layers according to the voidage of the dilated convolution layers; the sensing view field of the haze monitoring image is expanded, it is ensured that local haze monitoring features can be effectively extracted, and meanwhile the capturing capacity of the network for long-distance features is enhanced. According to the method, the information extraction capability can be improved on the premise of not increasing the calculation amount; efficient fog removal can be realized, and meanwhile, more textures and detail information are reserved, so that image details are clearer, and the blurring feeling caused by fog shielding is reduced; and the problems of blurring distortion, image darkness and the like of the defogged image are solved, so that the restored image is more natural and more conforms to the visual perception of human eyes, and the image restoration quality is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically to a dehazing method for low-exposure surveillance images of AOD-Net based on the self-attention mechanism. Background Technique

[0002] At present, video surveillance systems are growing with the development of image processing technology. From a single analog system to the current digital system, there has been a huge transformation in the surveillance mode. Video surveillance systems are widely used in various fields such as security, transportation, and industry. In the actual use process, in a haze scene, problems such as low visibility, poor contrast, and blurred images in the captured images directly affect the accuracy of surveillance recognition. Therefore, proposing an effective and feasible dehazing method has important significance for practical applications.

[0003] The dehazing method based on the physical model mainly analyzes the physical principle of haze formation and restores the image through, such as the atmospheric scattering model. It removes the influence of haze by estimating the transmittance in the image, thereby restoring clarity. It has a good dehazing effect in scenes with relatively light haze or fixed haze characteristics. However, it is not flexible enough when dealing with dynamic scenes, resulting in an unsatisfactory dehazing effect. The dehazing method based on image enhancement includes methods such as histogram equalization, contrast stretching, and color transfer functions to achieve the dehazing effect. It has the characteristics of fast calculation speed, simple calculation, strong real-time performance, and applicability during the dehazing process. However, in the actual dehazing process, it has a weak ability to restore complex haze scenes or image details, resulting in image noise amplification or detail loss.

[0004] According to the above problems, AOD-Net proposes an end-to-end deep learning network model. Based on the atmospheric scattering model, by estimating parameters in the model, the error during the dehazing process is reduced. However, since AOD-Net directly outputs the dehazed image instead of independently estimating the transmittance and atmospheric light, it has limitations in dealing with detail restoration. At the same time, it also causes a decrease in the brightness and contrast of the image, especially in areas that are originally darker in the image, which is more obvious. And the network focuses on local features and lacks global consistency, resulting in the inability to fully restore the distant view part or the overall structure of the image during the dehazing process, and the dehazing effect appears inconsistent or unnatural. Therefore, a dehazing method for low-exposure surveillance images of AOD-Net based on the self-attention mechanism is proposed to solve the above problems. Summary of the Invention

[0005] (1) Technical Problems to be Solved

[0006] In view of the deficiencies of the prior art, the present invention provides a dehazing method for low-exposure surveillance images of AOD-Net based on the self-attention mechanism, which solves the problems raised in the above background technique.

[0007] (2) Technical Solution

[0008] The present invention specifically adopts the following technical solutions to achieve the above purposes:

[0009] A dehazing method for low-exposure surveillance images of AOD-Net based on self-attention mechanism, comprising the following steps:

[0010] S1: Design dilated convolutions to replace the standard convolutions in AOD-Net. There are 5 convolutional layers in the network. Utilize the characteristics of dilated convolutions to adjust the dilation rates of the 5 dilated convolutional layers respectively, expand the receptive field of the haze surveillance images, ensure that local haze surveillance features can be effectively extracted, and at the same time enhance the network's ability to capture distant features;

[0011] S2: Add a Swin Transformer module before the fifth dilated convolution layer, establish the dependence relationship between global haze features and local haze features through the shifted window-based self-attention mechanism, capture the context information of the haze surveillance images, and help the network model the distant pixels in the haze surveillance images;

[0012] S3: Perform feature concatenation on the replaced multiple dilated convolutional layers respectively to fuse feature information of different scales;

[0013] S4: Add an RRDNet module, analyze the relationship between illumination, reflectivity, and noise using the Retinex model, estimate the noise iteratively, suppress the problem of noise amplification in the haze surveillance images caused by low exposure, enhance the overall restoration effect of low-exposure images, and achieve enhancement of the brightness and clarity of the images.

[0014] Further, in S1, input the haze surveillance images into the AOD-Net network, replace the standard convolutions with dilated convolutions. The convolution sizes of different convolutional layers are 1×1, 3×3, 5×5, 7×7, 3×3 respectively, and the dilation rates dilation are set to 1, 2, 4, 6, 1 respectively. The dilated convolution formula is:

[0015]

[0016] where y(i,j) represents the value of the output feature map at position (i,j), x(·) represents the input feature map, w(·) represents the dilated convolution kernel of size K×L, r represents the dilation factor, K represents the size of the dilated convolution kernel in the horizontal direction, and L represents the size of the dilated convolution kernel in the vertical direction.

[0017] Furthermore, in S2, a Swin Transformer module is added before the fifth-layer dilated convolution. Through the Swin Transformer module, the image is divided into patches of size 4×4, the embedding dimension is set to 96, and the number of attention heads is 3. The features of multiple adjacent patches are concatenated through patch merging operations, and then the dimension is reduced through a linear layer to gradually construct a hierarchical feature map;

[0018] Based on the shifted window strategy, that is, two different window partitioning methods are alternately used in consecutive modules: the first layer uses the conventional window partitioning (W-MSA), that is, starting from the upper left corner of the image, the image is evenly divided into non-overlapping windows; the next layer then translates the windows (SW-MSA), and the translation distance is In this way, the newly partitioned windows will cross the window boundaries of the previous layer to achieve cross-window information interaction; for two consecutive layers of modules, the calculation formula is:

[0019]

[0020]

[0021] Among them, represents the output features of the SW-MSA and S-MSA modules, z l represents the output features of the MLP module, LN represents layer normalization, W-MSA represents the self-attention module calculated under the conventional window, SW-MSA represents the self-attention module under the shifted window, and MLP represents a two-layer fully connected network;

[0022] The position information uses relative position bias, and the formula is:

[0023]

[0024] Among them, Q, K, and V represent the query, key, and value matrices respectively, d represents the dimension of the query / key, and B represents the relative position bias matrix,

[0025] Since the relative position within the window ranges from [-M+1, M-1] along each axis, a relatively small-sized bias matrix is parameterized as and the corresponding value in B is obtained by looking up the table according to the displacement of each position's effect.

[0026] Furthermore, in S3, a connection layer is used to perform multi-scale feature splicing on multiple dilated convolution layers respectively. The formulas for the 3 connection layers are:

[0027] F concat1 =Concat(F dconv1 ,Fdconv2 )

[0028] F concat2 = Concat(F dconv2 , F dconv3 )

[0029] F concat3 = Concat(F dconv1 , F dconv2 , F dconv3 , F dconv4 )

[0030] Among them, F dconv1 , F dconv2 , F dconv3 , F dconv4 respectively represent the feature maps processed by the dilated convolutional layers 1, 2, 3, and 4, and Concat() represents the operation of feature concatenation in the channel dimension.

[0031] Furthermore, in S4, the Retinex model is used to decompose the image into three parts: illumination, reflectance, and noise, and the noise is estimated and the illumination is restored by iteratively optimizing the loss function. The specific formula is:

[0032] I(x) = R(x)·S(x) + N(x)

[0033] Among them, I(x) represents the image generated in the previous step as the input image, R(x) represents the reflectance part of the image, S(x) represents the illumination part, and N(x) represents the noise part;

[0034] The reflectance branch and the illumination analysis use the sigmoid activation function to ensure that the output value is within the range of [0, 1], and the noise branch uses the tanh activation function to enable the network to adaptively separate the three parts;

[0035] Then, the estimated illumination component is adjusted by Gamma transformation. The formula is:

[0036]

[0037] Among them, S(x) γ represents the illumination component obtained in the decomposition stage, represents the adjusted illumination, and γ represents a predefined parameter used to control the overall brightness adjustment degree;

[0038] To remove the influence of noise, calculate the reflectance component after noise removal:

[0039]

[0040] Among them, The reflectivity component after denoising is denoted as, \(I(x)\) represents the input image, \(S(x)\) represents the illumination part, and \(N(x)\) represents the noise part;

[0041] The restored image is obtained by combining the adjusted illumination and the reflectivity component after noise removal:

[0042]

[0043] where, represents the finally output restored image, represents the reflectivity component after noise removal, represents the adjusted illumination component after Gamma transformation;

[0044] The quality decomposition and restoration are improved by iteratively minimizing the loss function. A comprehensive loss function \(L\) is designed based on the Retinex reconstruction loss, texture enhancement loss, and illumination-guided noise estimation loss.

[0045] Comprehensive loss formula:

[0046] \(L = L_{\text{Retinex}} + \lambda_1L_{\text{enhance}} + \lambda_2L_{\text{noise}}\) r +\(\lambda_1\) t \(L_{\text{enhance}}\) t +\(\lambda_2\) n \(L_{\text{noise}}\) n

[0047] where, \(L_{\text{Retinex}}\) r represents the Retinex reconstruction loss function, \(L_{\text{enhance}}\) t represents the enhancement loss function, \(L_{\text{noise}}\) n represents the illumination-guided noise estimation loss function, \(\lambda_1\) t and \(\lambda_2\) n represent the corresponding weight coefficients.

[0048] (III) Advantageous Effects

[0049] Compared with the prior art, the present invention provides a dehazing method for low-exposure surveillance images based on the self-attention mechanism, having the following advantageous effects:

[0050] By using dilated convolution to replace standard convolution, the present invention expands the receptive field of the convolutional layer to enhance the model's ability to capture large-scale context, and fuses feature images of different scales to effectively extract and fuse local features and long-distance features, improving the information extraction ability without increasing the computational complexity.

[0051] By adding the Swin Transformer module, when dealing with long-range dependencies and high-resolution images, the present invention extracts the global haze features and local haze feature information of the image through the shifted window self-attention mechanism. The output information contains rich global context information, improves the modeling of distant pixels in the image, realizes efficient fog removal, and at the same time retains more texture and detail information, making the image details clearer and reducing the blurring caused by fog occlusion.

[0052] By adding the RRDNet module, the present invention decomposes the image into illumination, reflectance, and noise, and realizes denoising by iteratively and effectively estimating the noise, solving problems such as blurred distortion and dark image appearance after fog removal, making the restored image more natural and more in line with the visual perception of the human eye, and further improving the quality of image restoration. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a schematic flowchart of the method of the present invention;

[0054] Figure 2 is a schematic diagram of the improved AOD-Net dehazing model of the embodiment of the present invention;

[0055] Figure 3 is a schematic diagram of the Swin Transformer module structure of the embodiment of the present invention;

[0056] Figure 4 is a schematic diagram of the shifted window strategy of the embodiment of the present invention;

[0057] Figure 5 is a schematic diagram of the RRDNet model of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0059] Embodiment

[0060] As Figures 1-5 shown, a method for dehazing low-exposure surveillance images based on self-attention mechanism proposed in an embodiment of the present invention includes the following steps:

[0061] S1: Design dilated convolutions to replace the standard convolutions in AOD-Net. There are 5 convolutional layers in the network. Utilize the characteristics of dilated convolutions to adjust the dilation rates of the 5 dilated convolutional layers respectively, expand the receptive field of haze monitoring images, ensure that local haze monitoring features can be effectively extracted, and at the same time enhance the network's ability to capture distant features;

[0062] As Figure 2 shown, input the monitored hazy image into the model. Estimate the value of K through the K value estimation module, and then in the image generation module, generate a clear image based on the atmospheric scattering model formula.

[0063] Atmospheric scattering model formula:

[0064] I(x) = J(x) × t(x) + A × [1 - t(x)]

[0065]

[0066] Among them, x represents the position of the image pixel. J(x) × t(x) represents the attenuated part of the incident light. A × [1 - t(x)] represents the part of the atmospheric light imaging. I(x) represents the pixel value of the input hazy monitoring image at position x. J(x) represents the pixel value of the clear haze-free monitoring image at position x before being scattered by atmospheric particles. A represents the global atmospheric light value. t(x) represents the transmittance at x.

[0067] Transmittance formula:

[0068] t(x) = e -βd(x)

[0069] Among them, β represents the atmospheric scattering coefficient, and d(x) represents the scene depth.

[0070] Integrate the two unknown variables A and t(x) into one unknown quantity K(x).

[0071] The formula for the value of K(x) is:

[0072]

[0073] Among them, I(x) represents the pixel value of the input hazy monitoring image at position x. A represents the global atmospheric light value. t(x) represents the transmittance at x. b represents a constant bias.

[0074] After arrangement, the formula is:

[0075] J (x) = k (x) I (x) - k (x) + b

[0076] Among them, J(x) represents the pixel value of the clear fog-free monitoring image before being scattered by atmospheric particles at position x, I(x) represents the pixel value of the input foggy monitoring image at position x, and b represents a constant bias.

[0077] In the K-value module, dilated convolution is used to replace the standard convolution, and the sizes of different convolution layers are set as: 1×1, 3×3, 5×5, 7×7, 3×3; the formula for the standard convolution is:

[0078]

[0079] Among them, y(i,j) represents the value of the output feature map at position (i,j), x(·) represents the feature map, w(·) represents the dilated convolution kernel of size K×L, K represents the size of the dilated convolution kernel in the horizontal direction, and K represents the size of the dilated convolution kernel in the vertical direction.

[0080] By introducing the dilation factor r on the basis of the standard convolution, the receptive field size is flexibly controlled by adjusting the dilation rate r. The formula for the dilated convolution is:

[0081]

[0082] Among them, y(i,j) represents the value of the output feature map at position (i,j), x(·) represents the feature map, w(·) represents the dilated convolution kernel of size K×L, K represents the size of the dilated convolution kernel in the horizontal direction, L represents the size of the dilated convolution kernel in the vertical direction, and r represents the dilation factor.

[0083] The dilation rates of the dilated convolution are set as dilation = 1, 2, 4, 6, 1 respectively.

[0084] S2: Add a Swin Transformer module before the fifth-layer dilated convolution, establish the dependence relationship between the global haze features and the local haze features through the shifted windowed self-attention mechanism, capture the context information of the haze monitoring image, and help the network model the pixels at a long distance in the haze monitoring image;

[0085] The specific content is as follows:

[0086] Ⅰ: First, divide the image into non-overlapping patches of size 4×4, set the window size to 7, the embedding dimension to 96, the number of attention heads to 3, and each patch is regarded as a tocken. Then, map the RGB pixels to a high-dimensional space through a linear embedding layer to obtain the initial token features. As the network deepens, adjacent multiple patch features are concatenated through the patch merging operation, and then the dimension is reduced through a linear layer to gradually construct a hierarchical feature map, enabling the network to capture multi-scale information from low-level to high-level, such as Figure 3As shown

[0087] II: Global self-attention calculates the relationship between each token and all other tokens, and the computational complexity grows quadratically with the number of tokens. Swin Transformer calculates self-attention within a fixed-size local window, solving the bottleneck in high-resolution image processing. The complexity formula for global self-attention is:

[0088] Ω(MSA) = 4hwC 2 + 2(hw) 2 C

[0089] where h and w represent the number of patches in the height and width directions respectively, and C represents the number of channels.

[0090] The complexity formula for window self-attention is:

[0091] Ω(W-MSA) = 4hwC 2 + 2M 2 hwC

[0092] where h and w represent the number of patches in the height and width directions respectively, C represents the number of channels, and M represents the number of patches in each window, with a size of 7.

[0093] This makes the overall complexity grow linearly with respect to h*w, greatly reducing the computational burden. Through the self-attention of local windows, while ensuring the ability to capture local features, the amount of computation is effectively controlled, enabling the model to process large-size, high-resolution images.

[0094] III: As Figure 4 shown, a shifted window strategy is adopted, that is, two different window partitioning methods are alternately used in consecutive modules: the first layer uses the conventional window partitioning (W-MSA), that is, starting from the upper left corner of the image, the image is evenly divided into non-overlapping windows; the next layer then translates the windows (SW-MSA) by a distance of (i.e., half of the window size), so that the newly partitioned windows will cross the window boundaries of the previous layer, realizing cross-window information interaction. For two consecutive layers of modules, the calculation formula is:

[0095]

[0096] where represents the output features of the SW-MSA and S-MSA modules, z l represents the output features of the MLP module, z l-1 represents the output of the previous module, and when l = 1, it represents the initial output. Represents the intermediate result output by the next module's self-attention module, z l+1 Represents the final output result after the (l + 1)-th module passes through the MLP module. LN represents layer normalization, W-MSA represents the self-attention module calculated under a regular window, SW-MSA represents the self-attention module under a shifted window, and MLP represents a two-layer fully connected network.

[0097] Ⅳ: Introduce position information and adopt relative position bias. The formula is:

[0098]

[0099] Among them, Q, K, and V represent the query, key, and value matrices respectively, d represents the dimension of the query / key, and B represents the relative position bias matrix.

[0100] Since the relative positions within the window range from [-M + 1, M - 1] along each axis, parameterize a bias matrix of a smaller size as And look up the corresponding value in B according to the displacement of the effect of each position.

[0101] S3: Perform feature concatenation on the replaced dilated convolutional layers respectively to fuse feature information of different scales. The feature concatenation formulas for the 3 connection layers are:

[0102] F concat1 = Concat(F dconv1 , F dconv2 )

[0103] F concat2 = Concat(F dconv2 , F dconv3 )

[0104] F concat3 = Concat(F dconv1 , F dconv2 , F dconv3 , F dconv4 )

[0105] Among them, F dconv1 , F dconv2 , F dconv3 , F dconv4 represent the feature maps processed by the dilated convolutional layers 1, 2, 3, and 4 respectively, and Concat() represents the feature concatenation operation in the channel dimension.

[0106] S4: Add the RRDNet module, use the Retinex model to analyze the relationships among illumination, reflectance, and noise, estimate the noise through iteration, suppress the problem of noise amplification in haze monitoring images caused by low exposure, enhance the overall restoration effect of low-exposure images, and achieve the enhancement of the brightness and clarity of the images.

[0107] The specific content is as follows:

[0108] Ⅰ: As Figure 5 shown, use the Retinex model to decompose the image into three parts: illumination, reflectance, and noise, and estimate the noise and restore the illumination by iteratively optimizing the loss function.

[0109] The formula of the Retinex model is:

[0110] I(x) = R(x)·S(x) + N(x)

[0111] Among them, I(x) represents the image generated in the previous step as the input image, R(x) represents the reflectance part of the image, S(x) represents the illumination part, and N(x) represents the noise part.

[0112] For the reflectance branch and illumination analysis, use the sigmoid activation function to ensure that the output value is within the range of [0, 1]. For the noise branch, use the tanh activation function to enable the network to adaptively separate the three parts.

[0113] Ⅱ: After obtaining the image decomposition result, perform corresponding processing on each component, and finally reconstruct the restored image. First, perform Gamma transformation adjustment on the estimated illumination component. The formula is:

[0114]

[0115] Among them, S(x) γ represents the illumination component obtained in the decomposition stage, represents the adjusted illumination, and γ represents a predefined parameter used to control the overall brightness adjustment degree.

[0116] Then, remove the influence of noise and calculate the reflectance component after noise removal:

[0117]

[0118] Among them, represents the reflectance component after denoising, I(x) represents the input image, S(x) represents the illumination part, and N(x) represents the noise part.

[0119] Finally, the restored image is obtained by combining the adjusted illumination and the reflectance component after noise removal:

[0120]

[0121] Among them, represents the restored image of the final output, represents the reflectivity component after noise removal, represents the adjusted illumination component after Gamma transformation.

[0122] Ⅲ: Improve quality decomposition and restoration through an iterative loss function, and design a comprehensive loss function L based on the Retinex reconstruction loss, texture enhancement loss, and illumination-guided noise estimation loss.

[0123] Comprehensive loss formula:

[0124] L = L r + λ t L t + λ n L n

[0125] Among them, L r represents the Retinex reconstruction loss function, L t represents the texture enhancement loss function, L n represents the illumination-guided noise estimation loss function, λ t and λ n represent the corresponding weight coefficients.

[0126] Retinex reconstruction loss function formula:

[0127]

[0128] Among them, I represents the input reconstructed image, R represents the reflectance component of the image, S represents the illumination component of the image, N represents the noise component of the image, S0(x) = max c∈{R,G,B} I c (x) represents the preliminary estimate of the illumination, and the L1 norm is used to ensure that the generated components can accurately reconstruct the input image.

[0129] Texture enhancement loss function formula:

[0130]

[0131] Among them, x and y represent the horizontal and vertical directions, w x and w y represent the weight factors, S represents the illumination component of the image, and represent the gradients of the illumination component S in the horizontal and vertical directions.

[0132] Formula for the weight w x :

[0133]

[0134] Among them, I g represents the grayscale version of the input image, represents the gradient of the grayscale version of the input image at coordinate x, G represents the Gaussian filter, represents the convolution operation. The weight w y formula is the same as that of w x in principle.

[0135] The weight w y formula:

[0136]

[0137] Illumination-guided noise estimation loss formula:

[0138]

[0139] Among them, N represents the noise component of the image, represents the hyperparameter used to balance the relative importance between noise amplitude constraints, R represents the reflection component of the image, and respectively represent the gradients of the reflection component R in the horizontal and vertical directions, w n and w r illumination-guided weights, and the formula is:

[0140] w n (x) = I(x)

[0141]

[0142] Among them, I(x) represents the input image, and represent the gradients of the reflection component R in the horizontal and vertical directions, and normalize() represents the normalization operation.

[0143] In this example, the training model will be trained in the environment of an Intel(R) Core TM i7-10700K processor + NVIDIA GeForce RTX4080Ti + Pytorch 2.3.1. The dataset will be divided into a training set, a validation set, and a test set according to the ratio of 8:1:1. Finally, the trained network model will be used to perform defogging operations on the monitoring images.

[0144] Aiming at the problems of feature information loss, poor detail restoration effect and obvious brightness reduction after defogging in the AOD-Net defogging process, the present invention uses dilated convolution to replace the standard convolution, adds the Swin Transformer module and the RRDNet module to improve the AOD-Net defogging network, and constructs an AOD-Net low-exposure monitoring image defogging method based on the self-attention mechanism. The proposed defogging method of shifted windowed self-attention mechanism multi-feature fusion can effectively reduce the blurring caused by fog occlusion, retain more texture and detail information, and solve the problems of blurred distortion and image darkening during the defogging process. Therefore, the method proposed in this paper can be used as an effective and feasible way to defog monitoring images, so that the restored images are more natural, more in line with the human visual perception, and further improve the quality of image restoration.

[0145] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An AOD-Net low-exposure monitoring image defogging method based on the self-attention mechanism, characterized in that: It includes the following steps: S1: Design dilated convolutions to replace the standard convolutions in the AOD-Net. There are 5 convolutional layers in the network. Using the characteristics of dilated convolutions, adjust the dilation rates of the 5 dilated convolutional layers respectively to expand the receptive field of the haze monitoring image, ensure that local haze monitoring features can be effectively extracted, and at the same time enhance the network's ability to capture distant features; S2: Add a Swin Transformer module before the fifth dilated convolution layer. Establish the dependence relationship between global haze features and local haze features through the shifted window self-attention mechanism, capture the context information of the haze monitoring image, and help the network model distant pixels in the haze monitoring image; S3: Perform feature concatenation on the replaced multiple dilated convolutional layers respectively to fuse feature information of different scales; S4: Add an RRDNet module. Use the Retinex model to analyze the relationship between illumination, reflectance, and noise. Estimate the noise iteratively to suppress the problem of noise amplification in the haze monitoring image caused by low exposure, enhance the overall restoration effect of low-exposure images, and achieve enhanced brightness and clarity of the image.

2. The AOD-Net low-exposure monitoring image dehazing method based on the self-attention mechanism according to claim 1, characterized in that: In the above S1, input the haze monitoring image into the AOD-Net network, replace the standard convolution with dilated convolution. The convolution sizes of different convolutional layers are 1×1, 3×3, 5×5, 7×7, 3×3 respectively, and set the dilation rates dilation to be 1, 2, 4, 6, 1 respectively. The dilated convolution formula is: Among them, y(i,j) represents the value of the output feature map at position (i,j), x(·) represents the input feature map, w(·) represents the dilated convolution kernel of size K×L, r represents the dilation factor, K represents the size of the dilated convolution kernel in the horizontal direction, and L represents the size of the dilated convolution kernel in the vertical direction.

3. A dehazing method for low-exposure surveillance images of AOD-Net based on self-attention mechanism according to claim 1, characterized in that: In the above S2, add a Swin Transformer module before the fifth dilated convolution layer. Through the Swin Transformer module, divide the image into patches of size 4×4, set the embedding dimension to 96, the number of attention heads to 3, concatenate the features of adjacent multiple patches through patch merging operation, and then reduce the dimension through a linear layer to gradually construct a hierarchical feature map; Based on the shifted window strategy, that is, two different window partitioning methods are alternately used in consecutive modules: the first layer uses the conventional window partitioning (W-MSA), that is, starting from the upper left corner of the image, the image is evenly divided into non-overlapping windows; the next layer then translates the windows (SW-MSA) by a distance of In this way, the newly partitioned windows will cross the window boundaries of the previous layer to achieve cross-window information interaction; for two consecutive layers of modules, the calculation formula is: Among them, represents the output features of the SW-MSA and S-MSA modules, and z l represents the output features of the MLP module, LN represents layer normalization, W-MSA represents the self-attention module calculated under a regular window, SW-MSA represents the self-attention module under a shifted window, and MLP represents a two-layer fully connected network; The position information adopts relative position bias, and the formula is: Among them, Q, K, and V represent the query, key, and value matrices respectively, d represents the dimension of the query / key, and B represents the relative position bias matrix. Since the value range of the relative position within the window along each axis is [-M + 1, M - 1], a bias matrix with a smaller size is parameterized as And the corresponding value in B is obtained by looking up the table for the displacement according to the effect of each position.

4. A method for dehazing low-exposure monitoring images of AOD-Net based on self-attention mechanism according to claim 2, characterized in that: In the above S3, use a connection layer to perform multi-scale feature concatenation on multiple dilated convolutional layers respectively. The formulas of the 3 connection layers are: F concat1 = Concat(F dconv1 , F dconv2 ) F concat2 = Concat(F dconv2 , F dconv3 ) F concat3 = Concat(F dconv1 , F dconv2 , F dconv3 , F dconv4 ) Among them, F dconv1 , F dconv2 , F dconv3 , F dconv4 respectively represent the feature maps processed by the dilated convolutional layers 1, 2, 3, and 4, and Concat( ) represents the feature concatenation operation in the channel dimension.

5. A method for dehazing low-exposure surveillance images of AOD-Net based on self-attention mechanism according to claim 1, characterized in that: In the above S4, use the Retinex model to decompose the image into three parts: illumination, reflectance, and noise, and estimate the noise to restore the illumination by iteratively optimizing the loss function. The specific formula is: I(x) = R(x)·S(x) + N(x) Among them, I(x) represents the image generated in the previous step as the input image, R(x) represents the reflectance part of the image, S(x) represents the illumination part, and N(x) represents the noise part; The reflectance branch and the illumination analysis use the sigmoid activation function to ensure that the output value is within the range of [0,1]. The noise branch uses the tanh activation function to enable the network to adaptively separate the three parts; Then, perform Gamma transformation adjustment on the estimated illumination component, and the formula is: Among them, S(x) γ represents the illumination component obtained in the decomposition stage, represents the adjusted illumination, and γ represents a predefined parameter used to control the overall brightness adjustment degree; Remove the influence of noise and calculate the reflectance component after noise removal: Among them, represents the denoised reflectance component, I(x) represents the input image, S(x) represents the illumination part, and N(x) represents the noise part; The restored image is obtained by combining the adjusted illumination and the reflectance component after noise removal: Among them, represents the restored image of the final output, represents the reflectivity component after noise removal, represents the adjusted illumination component after Gamma transformation; Improve the quality decomposition and restoration through an iterative loss function. Design a comprehensive loss function L based on the Retinex reconstruction loss, texture enhancement loss, and illumination-guided noise estimation loss; Comprehensive loss formula: L = L r + λ t L t + λ n L n Among them, L r represents the Retinex reconstruction loss function, L t represents the enhancement loss function, L n represents the illumination-guided noise estimation loss function, λ t and λ n represent the corresponding weight coefficients.