Haze image restoration method based on semi-supervised learning and dynamic perception attention U-shaped network
By adopting semi-supervised learning and dynamic perceptual attention U-shaped network in haze image recovery, combined with supervised and unsupervised learning branches, the problem of poor generalization ability of the model is solved, and more efficient haze image defog removal effect and visual quality are achieved.
Patent Information
- Application Number
- CN202411968665.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-30
AI Technical Summary
In the prior art, data sets are acquired through synthesis and have a certain domain spacing from natural paired data sets, resulting in poor generalization capabilities of the model.
The haze image recovery method based on semi-supervised learning and dynamic perceptual attention U-type network is adopted. By building supervised learning branches and unsupervised learning branches, the network is trained together with synthetic and natural data sets to improve the generalization ability of the model.
Through semi-supervised learning strategies and dynamic perceptual attention U-shaped network, natural and synthetic data sets can be effectively utilized, the performance of defogging networks in practical applications can be improved, the noise and color distortion of images after defogging are reduced, and images with excellent visual quality can be restored.
Smart Images

Figure CN120070259A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and particularly to a haze image restoration method based on semi-supervised learning and dynamic perception attention U-shaped network. Background Art
[0002] The importance of image dehazing extends to various fields, including remote sensing, surveillance, and autonomous driving. In remote sensing, dehazing technology can extract key information from satellite or aerial images, thus facilitating accurate analysis of the Earth's surface. In surveillance systems, dehazing plays a key role in enhancing the visibility of outdoor environments, thereby improving object detection and tracking capabilities. For autonomous driving, dehazing can ensure a clearer field of view for in-vehicle cameras, thus enhancing safety and reliability in complex driving scenarios.
[0003] The development of image dehazing algorithms can be traced back to early image enhancement techniques. Early dehazing techniques mainly relied on image enhancement methods such as histogram equalization and Retinex theory. These methods enhanced the visual effects of haze images by improving image contrast and brightness. Image enhancement methods are simple to implement and computationally efficient, but the disadvantage is that they cannot effectively restore the true details of images, and the effect is limited when dealing with severe haze. With a deeper understanding of the causes of haze, researchers began to adopt physical models such as the atmospheric scattering model and combine image prior knowledge for image restoration. These methods restore haze-free images by estimating atmospheric light and transmittance. Representative algorithms include the dark channel prior. The advantage of this type of algorithm is that it can better restore image details, but it has high requirements for the estimation of physical model parameters and may be distorted in complex scenarios.
[0004] With the booming development of deep learning technology, researchers began to use convolutional neural networks (CNNs) to directly restore clear images from haze images. They can either estimate parameters based on the atmospheric scattering model or directly perform end-to-end learning without using a physical model. Such methods can automatically learn complex feature representations, thus improving the dehazing effect, but they require a large number of paired hazy images and clear images for training. However, most current datasets are obtained synthetically, having a certain domain gap with natural paired datasets, resulting in poor model generalization ability and still needing further improvement. Summary of the Invention
[0005] This application provides a haze image restoration method based on semi-supervised learning and dynamic perception attention U-shaped network, which can be used to solve the technical problem that the dataset is obtained synthetically, having a certain domain gap with natural paired datasets, resulting in poor model generalization ability.
[0006] A haze image restoration method based on semi-supervised learning and dynamic perception attention U-shaped network, the method includes:
[0007] Step A: Construct the overall architecture of the defogging network, namely the dynamic perception attention U-shaped network. The U-shaped network is divided into an encoding part, a feature transformation part, and a decoding part from the input end to the output end, and can output a fog-free image end-to-end.
[0008] Step B: Construct a supervised learning branch to perform supervised training on the defogging network. Use synthetic haze images as the input of the defogging network, use real-world clean images as labels, and adopt mean square error loss, perceptual loss, and adversarial loss for supervised training.
[0009] Step C: Construct an unsupervised learning branch to perform unsupervised training on the defogging network. Use natural haze images as the input of the defogging network, and adopt dark channel loss and total variation loss for unsupervised training.
[0010] Step D: Use the dynamically perceived attention U-shaped network trained in Step B and Step C, input a haze image, and output a fog-free image end-to-end.
[0011] Compared with the prior art, the significant advantages of the present invention are as follows: 1. Use a semi-supervised learning strategy, which uses both labeled data and unlabeled data during the training process. This enables the defogging network to train using both natural datasets and synthetic datasets, and can not only learn the mapping relationship from foggy images to fog-free images through prior knowledge, but also learn the feature distribution of real-world fog-free images, improving the performance of the defogging network in practical applications. 2. Design a deformable convolutional residual block, which enhances the feature extraction network based on deformable convolution and residual learning, improving the network's ability to model geometric transformations and enabling it to extract more effective feature representations. 3. Extract and analyze channel-specific spatial information through content-guided attention to generate a channel-level spatial importance distribution map, effectively coping with the non-uniformity of haze distribution. 4. Propose a hierarchical feature fusion module to replace simple skip connections, effectively fusing shallow and deep features, enhancing the information flow from shallow to deep, and facilitating the restoration of clearer images. Brief Description of the Drawings
[0012] Figure 1 It is a flowchart for haze image restoration in the actual work provided by this application.
[0013] Figure 2 It is an overall architecture diagram of a haze image restoration method based on semi-supervised learning and a dynamically perceived attention U-shaped network designed by the method provided by this application.
[0014] Figure 3 It is a structural diagram of a dynamically perceived attention U-shaped network designed by the method provided by this application.
[0015] Figure 4The structural diagram of the deformable convolutional residual block for the method provided in this application.
[0016] Figure 5 The structural diagram of the content-guided attention for the method provided in this application.
[0017] Figure 6 The structural diagram of the hierarchical feature fusion module for the method provided in this application. Detailed implementation manners
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the implementation manners of this application in detail with reference to the accompanying drawings.
[0019] First, the embodiments of this application will be introduced below with reference to the accompanying drawings.
[0020] Combined with Figure 1 , the present invention relates to a haze image restoration method based on semi-supervised learning and a dynamic perception attention U-shaped network, which mainly includes: Step A, constructing a dehazing network - the overall architecture of the dynamic perception attention U-shaped network; the U-shaped network is divided into an encoding part, a feature transformation part, a decoding part, and a feature fusion part from the input end to the output end, and can obtain a haze-free image end-to-end; Step B, constructing a supervised learning branch; using the synthesized haze image as the input of the U-shaped network, using the clean image in the real world as the label, and adopting the mean square error loss, perceptual loss, and adversarial loss for supervised training; Step C, constructing an unsupervised learning branch; using the natural haze image as the input of the U-shaped network, and adopting the dark channel loss and total variation loss for unsupervised training, and sharing the weights with the network parameters obtained in Step B during the training process. Step D, using the trained model parameters to construct a dynamic perception attention U-shaped network, inputting the haze image, and outputting the haze-free image restored by the dehazing network.
[0021] The network architecture diagram of the haze image restoration method based on semi-supervised learning and dynamic perception attention U-shaped network is as Figure 2 shown, and the method includes the following steps:
[0022] Step A, constructing a dehazing network, that is, the overall architecture of the dynamic perception attention U-shaped network; the U-shaped network is divided into an encoding part, a feature transformation part, and a decoding part from the input end to the output end, and can obtain a haze-free image end-to-end;
[0023] Step B, constructing a supervised learning branch to perform supervised training on the dehazing network constructed in Step A; using the synthesized haze image as the input of the dehazing network constructed in Step A, using the clean image in the real world as the label, and adopting the mean square error loss, perceptual loss, and adversarial loss for supervised training;
[0024] Step C: Construct an unsupervised learning branch to perform unsupervised training on the dehazing network constructed in Step A. Use natural haze images as network inputs, and adopt dark channel loss and total variation loss for unsupervised training.
[0025] Step D: Utilize the dynamically aware attention U-shaped network that has been trained in Steps B and C. Input a haze image and output a haze-free image end-to-end.
[0026] Step A: Construct the overall architecture of the dehazing network - the dynamically aware attention U-shaped network. The overall architecture of the dynamically aware attention U-shaped network is as Figure 3 shown. The U-shaped network is divided into an encoding part, a feature transformation part, and a decoding part from the input end to the output end, and can obtain a haze-free image end-to-end. It includes:
[0027] Step A1: Construct the encoder of the U-shaped network to obtain multi-semantic feature layers. Among them, the encoder uses a 3×3 convolutional layer to adjust the number of channels, uses a deformable convolutional residual block to dynamically extract the features of the input image, and uses a 3×3 convolutional layer with a stride of 2 for downsampling. The deformable convolutional residual block and the downsampling layer are repeated twice, and a total of three feature layers with different scales are obtained.
[0028] Among them, the core part of the deformable convolutional residual block is the deformable convolution, and its formula is as follows:
[0029]
[0030] where x is the input feature, p is the center position of the convolutional sampling, p k is the offset of the center point, p k ∈{(-1,-1),(-1,0),…,(0,1),(1,1)}. Δp k and Δm k are the offsets and weight coefficients of K sampling points respectively, and the value range of Δm k is [0,1]. Δp k is a real number with an unconstrained range, resulting in p + p k + Δp k being a real number. When calculating x(p + p k + Δp k ), the bilinear interpolation method will be applied to confirm the sampling points on the original feature map. Δp k and Δm k are both learnable parameters, which are obtained through two separate convolutional layers.
[0031] Furthermore, a Deformable Convolutional Residual Block (DCRB) is constructed using deformable convolutions. The structure diagram of the Deformable Convolutional Residual Block is as shown in Figure 4 . One deformable convolution and a 3×3 convolution block are used. After each convolution, a ReLU activation layer is added, and two residual connections are introduced to directly connect the input feature map with the convolved feature map. The formula for the Deformable Convolutional Residual Block is as follows:
[0032]
[0033] where max(0, x) represents the ReLU activation function, represents a convolutional layer with a kernel size of 3, and during convolution, padding is performed in the same way at the edges to ensure that the size of the output feature layer remains unchanged. DeConv(x) represents the deformable convolution block.
[0034] Step A2: Construct the feature transformation part. Based on the minimum-size feature layer obtained in the encoding process of Step A1, feature transformation is performed to extract multi-channel semantic information of the low-resolution feature layer. The feature transformation part includes a series of deformable convolutional residual blocks and a hierarchical feature fusion module. Among them, the Hierarchical Feature Fusion with Attention (HFFA) includes a content-guided attention mechanism and a 1×1 convolutional layer, which can mix the information of shallow and deep feature maps.
[0035] Furthermore, the core of the hierarchical feature fusion module is the Content Guided Attention (CGA) mechanism, which can extract a spatial importance distribution map specific to the feature channels. The structure diagram of the content-guided attention is shown in Figure 5 , and this process is expressed as the following formula:
[0036]
[0037] In formula (3), represents the channel attention map, is the spatial attention map. max(0, x) represents the ReLU activation function, represents a convolutional layer with a kernel size of k×k. represents global average pooling in the spatial dimension, and represent global average pooling and global max pooling in the channel dimension respectively. It means concatenating two tensors with the same size along the channel dimension.
[0038] In formula (4), X represents the input feature map, and W c +W s means adding the spatial attention map and the channel attention map to obtain a channel-specific hybrid attention map. Since the two have different sizes, the operation follows the broadcasting mechanism. CS(·) represents shuffling in the channel dimension. represents grouped convolution, which reduces the computational amount compared with ordinary convolution. σ represents the sigmoid activation function, and after normalization, the spatial importance distribution map W is obtained.
[0039] Furthermore, the structure of the hierarchical feature fusion module is as Figure 6 shown. Using the low-level feature layer after downsampling and the corresponding high-level feature layer in the encoding process as inputs, first add the two layers of features to mix the high-level semantic information and the low-level semantic information; then use content-guided attention to calculate the spatial importance map, and multiply it with the respective feature maps to obtain refined feature maps. Finally, add the two original feature maps and the two refined feature maps, and then pass through a 1×1 convolutional layer to obtain the final fused feature map. This process is expressed by the formula as follows:
[0040]
[0041] Among them, F fuse represents the fused feature map, F low and F high respectively represent the bottom layer feature and the corresponding high layer feature, W represents the spatial importance map obtained through content-guided attention, is a 1×1 convolutional layer.
[0042] Step A3: Construct the decoder of the U-shaped network, and gradually perform feature recovery and feature fusion based on the fused feature map obtained in step A2. Among them, the encoder uses deformable convolutional residual blocks to extract features, and uses a 3×3 transposed convolutional layer with a stride of 2 for upsampling. The upsampling and deformable convolutional residual blocks will be repeated twice. Similar to the encoding process, three feature layers corresponding to different scales are obtained. Finally, use a 3×3 convolutional layer to adjust the number of channels to 3 and output the haze-free image. Among them, before upsampling, the hierarchical feature fusion module is used to fuse the feature layer after downsampling in step A1 again.
[0043] Furthermore, in step B, construct a supervised learning branch to perform supervised training on the haze removal network constructed in step A; use the synthetic haze image as the input of the haze removal network constructed in step A, use the clean image in the real world as the label, and adopt mean square error loss, perceptual loss and adversarial loss for supervised training; including:
[0044] Step B1: Introduce the mean square error loss to measure the pixel difference between the natural fog-free image and the predicted fog-free image. The loss function is expressed as follows:
[0045]
[0046] where represents n b represents the number of labeled data in a batch, N(I i ) represents the fog-free image predicted after passing the synthetic foggy image I i through the U-shaped network constructed in Step A, represents the natural fog-free image corresponding to I i , and ‖‖ 2 represents the L2-norm.
[0047] Step B2: Introduce the adversarial loss, which can improve the clarity and edge details of the predicted image. The loss function is expressed as follows:
[0048]
[0049] where, represents the discriminator, which needs to be trained together with the U-shaped network constructed in Step A. The first term of this loss represents the probability that the real image is true when passing through the discriminator, and the second term represents the probability that the predicted image is false when passing through the discriminator. The optimization process of this loss is to deceive the generator and reduce the probability that it considers the real image to be true and the predicted image to be false.
[0050] Step B3: Introduce the perceptual loss, which can establish the connection between the defogging task and the high-level task and improve the visual effect of the image. The loss function is expressed as follows:
[0051]
[0052] where F represents the feature extractor of the pre-trained object detection network. In the present invention, the VGG-19 network pre-trained on the COCO dataset is used. and are the feature maps of the predicted image and the natural image after passing through the feature extractor respectively, and then their L2-norms are calculated.
[0053] Step B4: Construct the composite loss function loss s , use the synthetic haze image as the input and the clean image of the real world as the label to train the dynamic spatial perception U-shaped network. The composite loss function is as follows:
[0054] loss s =αL c +βL a +γL p (9)
[0055] Among them, {α, β, γ} are hyperparameters that control the weights of each loss. In the present invention, they are respectively set as {α = 1, β = 0.01, γ = 0.001}
[0056] Further, in step C, an unsupervised learning branch is constructed to perform unsupervised training on the dehazing network constructed in step A; natural haze images are used as the network input, and dark channel loss and total variation loss are used for unsupervised training. It includes:
[0057] In step C1, the dark channel loss is introduced. The dark channel prior can be used to help the predicted image obtain the same statistical characteristics as the natural image. The dark channel loss function is as follows:
[0058]
[0059] Formula (11) represents the calculation formula of the dark channel, where I c represents the three rgb color channels in I, and Ω(x) represents the local block centered on x. Formula (10) is the calculation of the dark channel loss, and ‖‖ 1 represents the L1-norm sparse prediction image of the dark channel.
[0060] In step C2, the total variation loss is introduced, which can help the predicted image retain the structure and details of the natural fog image. The total variation loss function is as follows, where and respectively represent the horizontal and vertical gradient operators.
[0061]
[0062] In step C3, a composite loss function loss u is constructed, and natural haze images are used as the input of the U-shaped network to train the dynamic spatial perception U-shaped network. The loss function is as follows:
[0063] loss u = εL D + δL T (13)
[0064] Among them, {ε, δ} are hyperparameters that control the weights of each loss. In the present invention, they are respectively set as {ε = 10 -5 , δ = 10 -5}.
[0065] Further, the training details are as follows. At the beginning of training, first randomly select n b labeled samples. After these samples are processed by the network, the supervised loss is calculated. Then, randomly select n bAn unlabeled sample, these samples are processed by the same network, and the unsupervised loss is calculated. Finally, using the calculated supervised loss and unsupervised loss, the network parameters of the supervised branch and the unsupervised branch are updated in turn through the backpropagation algorithm. It should be noted that the dehazing networks of the unsupervised branch and the supervised branch share network weights during training, and the convergence of the supervised branch is taken as the end point of training completion.
[0066] Furthermore, in step D, using the dynamically aware attention U-shaped network trained in steps B and C, input the hazy image, and output the haze-free image end-to-end.
[0067] In summary, the present invention innovatively constructs a dynamically aware attention U-shaped network by combining the dynamic feature extraction ability of deformable convolution and the perception ability of the attention mechanism for the unevenness of foggy images, significantly improving the expression ability of the dehazing model. Then, using the MSE loss, adversarial loss, and perceptual loss, the U-shaped network is trained on the synthetic dataset to form a supervised learning branch; subsequently, using the dark channel loss and total variation loss, the U-shaped network is further trained and fine-tuned on the natural foggy image dataset to construct an unsupervised learning branch; finally, the trained network is used to dehaze the foggy image. The present invention can effectively reduce the noise and color distortion of the dehazed image, and restore an image with excellent visual quality, which is applicable to multiple fields such as remote sensing, autonomous driving, and surveillance analysis, showing broad application prospects and practical application values.
[0068] The above-described embodiments of the present application do not constitute a limitation on the protection scope of the present application.
Claims
1. A haze image restoration method based on semi-supervised learning and dynamic perception attention U-type network, characterized in that: The method comprises: Step A, constructing the overall architecture of the defogging network, i.e., the dynamic perception attention U-type network; the U-type network is divided into an encoding part, a feature conversion part, and a decoding part from the input end to the output end, and can output a haze-free image end-to-end; Step B: construct a supervised learning branch to perform supervised training on the dehazing network; use synthetic haze images as the input of the dehazing network, use real-world clean images as labels, and adopt mean square error loss, perceptual loss, and adversarial loss for supervised training; Step C: construct an unsupervised learning branch to perform unsupervised training on the dehazing network; use natural haze images as the input of the dehazing network, and adopt dark channel loss and total variation loss for unsupervised training; In step D, the dynamic perception attention U-type network trained in steps B and C is used as input to generate a haze image and output a haze-free image end-to-end.
2. The method according to claim 1, characterized in that: Step A, constructing the overall architecture of the dehazing network, i.e., the dynamic perception attention U-type network; including: Step A1, constructing an encoder of a U-shaped network to obtain a multi-semantic feature layer; wherein the encoder uses a 3×3 convolutional layer to adjust the number of channels, uses a deformable convolutional residual block to dynamically extract input image features, and uses a 3×3 convolutional layer with a step size of 2 for downsampling; the deformable convolutional residual block and the downsampling layer are repeated twice, and a total of three feature layers of different scales are obtained; Among them, the core part of the deformable convolution residual block is the deformable convolution, and the formula is as follows: Among them, x is the input feature, p is the center position of the convolution sampling, and p k is the offset of the center point, p k ∈{(-1,-1),(-1,0),…,(0,1),(1,1)}; Δp k and Δm k are the offset and weight coefficient of K sampling points, Δm k The value range of is [0,1]; Δp k is a real number with an unconstrained range, resulting in p+p k +Δp k is a real number. k +Δp k ) Apply bilinear interpolation to confirm the sampling points on the original feature map; Δp k and Δm k are all learnable parameters obtained through two separate convolutional layers. Use deformable convolution to construct a deformable convolution residual block (DCRB); use a deformable convolution and a 3×3 convolution block, add a ReLU activation layer after each convolution layer, and introduce two residual connections to directly add the input feature map to the feature map after convolution; the formula of the deformable convolution residual block is as follows: Among them, max(0,x) represents the ReLU activation function, Denotes a convolutional layer with a kernel size of 3, and the edges are padded in the same way during convolution; DeConv(x) denotes a deformable convolutional block; Step A2, constructing a feature conversion part, performing feature conversion based on the minimum size feature layer obtained in the encoding process of step A1, and extracting multi-channel semantic information of the low-resolution feature layer; the feature conversion part includes a series of deformable convolution residual blocks and a hierarchical feature fusion module; wherein the hierarchical feature fusion module (Hierarchical Feature Fusion with Attention, HFFA) includes a content-guided attention mechanism and a 1×1 convolutional layer, which can mix shallow and deep feature information; The core of the hierarchical feature fusion module is the content guided attention (CGA) mechanism, which can extract the spatial importance distribution map specific to the feature channel. The formula is as follows: In formula (3), represents the channel attention map, is the spatial attention map; max(0,x) represents the ReLU activation function, represents a convolutional layer with a kernel size of k×k; Indicates global average pooling in the spatial dimension, and They represent global average pooling and global maximum pooling in the channel dimension respectively; Indicates concatenating two tensors of the same size in the channel dimension; In formula (4), X represents the input feature map, W c +W s It means adding the spatial attention map and the channel attention map to obtain a channel-specific mixed attention map. Since the sizes of the two are inconsistent, the broadcast mechanism is followed; CS(·) means shuffling in the channel dimension. represents grouped convolution; σ represents the sigmoid activation function, and after normalization, the spatial importance distribution map W is obtained; The structure of the hierarchical feature fusion module uses the low-level feature layer after downsampling in the encoding process and the high-level feature layer of the corresponding level as input, adds the two layers of features, and mixes the high-level semantic information with the low-level semantic information; then uses content-guided attention to calculate the spatial importance map, and then multiplies it with each feature map to obtain a refined feature map; finally, the two original feature maps and the two refined feature maps are added, and then passed through a 1×1 convolution layer to obtain the final fused feature map; the formula is as follows: Among them, F fuse represents the fusion feature map, F low and F high Represent the underlying features and the corresponding high-level features respectively, W represents the spatial importance map obtained by content-guided attention, It is a 1×1 convolution layer; Step A3, construct a decoder of the U-type network, and gradually perform feature recovery and feature fusion based on the fused feature map obtained in step A2; wherein, the encoder uses a deformable convolution residual block to extract features, uses a 3×3 deconvolution layer with a step size of 2 for upsampling, and repeats the upsampling and deformable convolution residual block twice to obtain three feature layers of corresponding scales; finally, a 3×3 convolution layer is used to adjust the number of channels to 3, and a haze-free image is output; wherein, before upsampling, the hierarchical feature fusion module is used to fuse the feature layer after downsampling in step A1 again.
3. The method according to claim 1, characterized in that Step B: construct a supervised learning branch to perform supervised training on the dehazing network constructed in step A. Use the synthetic haze image as the input of the dehazing network constructed in step A, use the clean image of the real world as the label, and adopt mean square error loss, perceptual loss and adversarial loss for supervised training. include: Step B1, introduce the mean square error loss, the mean square error loss function is expressed as: Among them, represents n b Represents the number of labeled data in a batch, N(I i ) represents the synthetic fog image I i The haze-free image predicted by the U-shaped network constructed in step A, I i The corresponding natural haze-free image, ‖‖2 represents the L2-norm; Step B2, introduce adversarial loss, and the adversarial loss function is expressed as follows: in, represents the discriminator, which is trained together with the U-shaped network constructed in step A; the first term of the adversarial loss represents the probability that the real image is true when it passes through the discriminator, and the second term represents the probability that the predicted image is false when it passes through the discriminator; the optimization process of the adversarial loss is to deceive the generator and reduce the probability of believing that the real image is true and the predicted image is false; Step B3, introduce perceptual loss, and the perceptual loss function is expressed as follows: Among them, F represents the pre-trained target detection network feature extractor, using the VGG-19 network pre-trained on the COCO dataset, and The feature maps of the predicted image and natural image after the feature extractor are respectively calculated, and then their L2-norms are calculated; Step B4: Construct a composite loss function loss s , using synthetic haze images as input and real-world clean images as labels to train a dynamic spatially aware U-shaped network, the composite loss function is as follows: loss s =αL c +βL a +γL p (9) Among them, {α, β, γ} are hyperparameters that control the weight of each loss, and are set to {α=1, β=0.01, γ=0.001} respectively.
4. The method according to claim 1, characterized in that Step C, construct an unsupervised learning branch to perform unsupervised training on the dehazing network constructed in step A; Use natural haze images as network input, and use dark channel loss and total variation loss for unsupervised training. Including: Step C1, introduce dark channel loss, use dark channel prior to help predict the image to obtain the same statistical characteristics as natural images, the dark channel loss function is as follows: D(I)=min c∈{r,g,b} [my y∈Ω(x) IN c (y)] (11) Formula (11) represents the calculation formula of the dark channel, where I c represents the three color channels of RGB in I, Ω(x) represents the local block centered on x; formula (10) is the calculation of the dark channel loss, ‖‖1 represents the L1-norm, which can sparsely predict the dark channel of the image; Step C2, introduce the total variation loss, the total variation loss function is as follows, where, and Represent the horizontal and vertical gradient operators respectively: Step C3, construct the composite loss function loss u , using natural haze images as U-network input to train the dynamic space perception U-network, the composite loss function is as follows: loss u =εL D +δL T (13) Among them, {ε, δ} are hyperparameters that control the weight of each loss, which are set to {ε=10 -5 ,δ=10 -5 }.
5. The method according to claim 4, characterized in that At the beginning of training, we first randomly select n b After the labeled samples are processed by the network, the supervised loss is calculated; Then, randomly select n b unlabeled samples are processed by the same network and the unsupervised loss is calculated; Finally, the network parameters of the supervised and unsupervised branches are updated in sequence through the back-propagation algorithm using the calculated supervised and unsupervised losses. The dehazing networks of the unsupervised branch and the supervised branch share network weights during the training process, and the convergence of the supervised branch is regarded as the end point of the training.
Citation Information
Patent Citations
Image defogging method based on semi-supervision
CN114155165A
Defogging method based on foggy day traffic road image
CN118365558A
Cited By
Image defogging system and method for low-altitude scene
CN120598820A
Network training and application method and equipment based on spatial transformation and comparative learning
CN120852243A
Network training, application method and device based on spatial transformation and contrast learning
CN120852243B
Defogging adaptive method for synthesizing remote sensing image based on pseudo fog
CN121937323A