Power transmission line icing image defogging method based on improved CycleGAN

By introducing a channel attention mechanism in CycleGAN, the problem of loss of ice-covering feature details and incomplete fog removal in the traditional fog removal method is solved, and a higher quality fog removal effect of transmission line ice-covering image is achieved.

CN120070253APending Publication Date: 2025-05-30ELECTRIC POWER RES INST OF EAST INNER MONGOLIA ELECTRIC POWER +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510147916.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The traditional defog removal method has problems such as loss of ice feature details and incomplete defog removal when processing the ice-covered image of transmission lines.

Method used

The channel attention mechanism is introduced into the generator of CycleGAN, which enhances the generator's attention to the characteristics of the ice-covered section of the transmission line, and improves the loss function of the CycleGAN model to enhance the detail retention and feature extraction capabilities.

Benefits of technology

It effectively improves the defogging effect of the ice-covered image of the transmission line and enhances the clarity of the defogging image. Especially in complex weather conditions, it can better retain key information such as ice edges and ice-thick texture characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070253A_ABST
    Figure CN120070253A_ABST
Patent Text Reader

Abstract

The invention discloses a power transmission line icing image defogging method based on an improved CycleGAN, and is applied to the technical field of image processing. The method comprises the following steps: firstly, acquiring a foggy image and a fogless image in an ice-coated power transmission line, and preprocessing the images; and then, a channel attention mechanism is introduced into the generator to a residual block to adaptively adjust feature weights of different channels in the image, and the decoder part adopts a bilinear interpolation and deconvolution combined method to perform image restoration. A loss function of the generator is constructed by adopting a method of combining confrontation loss, cyclic consistency loss, characteristic consistency loss and a channel attention mechanism, and the detail retention and feature extraction capabilities of the defogging model are enhanced. And finally, outputting a complete defogged image through model training. According to the method, the haze noise in the icing image of the power transmission line under the complex meteorological condition can be effectively removed, the detail information in the image is reserved to the greatest extent, and the fog-free image with relatively high definition is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and more particularly to a method for defogging an ice-covered image of a power transmission line based on an improved CycleGAN. Background Art

[0002] The icing of transmission lines is an important problem faced by the power system. Icing can cause damage to transmission lines, and in severe cases, it can cause the power system to be unable to operate normally. The image monitoring system of transmission lines is usually used to monitor the icing of transmission lines in real time. However, under severe weather conditions, the monitoring images are often interfered by environmental factors such as haze and snowfall, resulting in a decrease in image quality and an inability to accurately identify the thickness of ice on the transmission lines and determine the icing status. Therefore, image defogging technology is an important means to improve the monitoring of icing on transmission lines.

[0003] With the development of deep learning, unsupervised learning models based on methods such as generative adversarial networks (GAN) have made significant progress in the field of image defogging. GAN can generate high-quality defogging images with better versatility and visual effects through adversarial training of generators and discriminators. Among them, the cyclic generative adversarial network (CycleGAN) is an unsupervised learning method that does not require paired training data, and is particularly suitable for scenes where paired foggy and fog-free image samples are not available. In the task of defogging of ice-covered transmission lines, detailed information such as the outline of the transmission line and the texture characteristics of the ice thickness are crucial to the judgment of image quality. However, CycleGAN still has shortcomings when processing images of complex scenes, especially in terms of detail retention. Therefore, how to provide a method for defogging ice-covered transmission line images based on an improved CycleGAN is a problem that technicians in this field urgently need to solve. Summary of the invention

[0004] In view of this, the present invention provides a method for defogging ice-covered images of power lines based on improved CycleGAN, which solves the problems of loss of ice feature details and incomplete defogging in traditional defogging methods. The channel attention mechanism is introduced into the generator of CycleGAN to improve the generator's attention to the features of ice-covered sections of power lines, enhance the detail retention and feature extraction capabilities of the defogging model, and help the CycleGAN defogging model better capture important details in ice-covered images such as cable details and ice thickness features, effectively improving the defogging effect of ice-covered images of power lines.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] A method for defogging ice-covered images of power transmission lines based on improved CycleGAN includes the following steps:

[0007] S1. Obtain the fog-free images and foggy images of the ice-covered transmission line, and preprocess the images, including image denoising, image normalization, and image cropping and adjustment;

[0008] S2. Construct the generator of the improved CycleGAN model, including an encoder, a residual block, and a decoder connected in sequence, and input the preprocessed foggy images into the generator for learning;

[0009] S3. Construct the discriminator of the improved CycleGAN model, and the discriminator judges the authenticity of the images generated by the generator and the real images;

[0010] S4. Use adversarial loss, cycle consistency loss, feature consistency loss, and channel attention mechanism loss to construct the loss function of the generator of the improved CycleGAN model;

[0011] S5. Use the foggy images of the ice-covered transmission line under harsh meteorological conditions and the fog-free images under normal weather conditions to train the generator and the discriminator. The generator gradually generates clearer de-fogged images through adversarial training and cycle consistency loss, and the discriminator judges the authenticity of the de-fogged images through adversarial training;

[0012] S6. After the improved CycleGAN model is trained, the generator generates de-fogged images according to the input foggy images.

[0013] Optionally, S1 is specifically as follows:

[0014] S11. Through the image acquisition device deployed on the transmission line, collect the fog-free images of the ice-covered transmission line under normal climate conditions and the foggy images under harsh meteorological conditions. The collected images include the transmission line itself, the surrounding environment, and the part covered by ice;

[0015] S12. Denoise the collected images through Gaussian filtering. Gaussian filtering smooths the images to reduce the noise introduced by environmental and equipment factors, making the overall details of the images more coherent;

[0016] S13. Perform pixel normalization processing on the collected images, divide each pixel value of the images by 255, and scale the pixel values of the images to the interval [0, 1] to ensure that the later de-fogging model can better process images with different brightnesses;

[0017] S14. Crop the images according to the image content, remove redundant background information, and crop the images to a size of 256×256 pixels.

[0018] Optionally, the image generation process of the generator is specifically as follows:

[0019] The encoder uses multiple convolutional layers to extract low-level features of the input image. In the convolutional layer, the input image is downsampled through convolution operations, gradually reducing the size of the feature map and extracting deep-level features. The convolutional layer contains multiple convolutional kernels that slide on the input feature map based on dot product operations to calculate the feature values at each position, thereby extracting specific spatial features;

[0020] The residual block is used to extract low-level and high-level features from the input image through the convolutional layer. The global average pooling and fully connected layer of the channel attention mechanism are used to calculate the weights of each channel, select the channels containing the key information of ice coverage, enhance their features, and perform residual addition on the weighted feature map and the original feature map to form a residual connection;

[0021] In the decoder, bilinear interpolation is used as the initial upsampling operation, and the feature map is further refined through deconvolution operations, enhancing the detail retention effect of defogging while ensuring the smoothness of upsampling.

[0022] Optionally, the image processing process of the encoder is specifically as follows:

[0023] The first layer of the encoder uses a convolutional layer to extract the initial features of the image, performs batch normalization on the output after convolution, and uses the ReLU activation function to introduce non-linearity and enhance the feature expression ability;

[0024] The second layer of the encoder uses a convolutional layer for downsampling to further extract features, reducing the size of the output feature map by half, performing batch normalization on the output after convolution, and using the ReLU activation function;

[0025] The third layer of the encoder uses a convolutional layer to further downsample the feature map, extract higher-level features, further reduce the resolution of the feature map, perform batch normalization on the output after convolution, and use the ReLU activation function.

[0026] Optionally, the image processing process of the residual block is specifically as follows:

[0027] Convolve the input feature map X, and then perform batch normalization and ReLU activation:

[0028] X 1 = ReLU(BN(Conv(X)))

[0029] Perform another convolution and batch normalization operation without using the activation function:

[0030] X 2 = BN(Conv(X 1 ))

[0031] In the formula, X 1is the feature map of the first convolution output, X 2 is the feature map of the second convolution output, BN represents batch normalization, and Conv represents convolution;

[0032] For the output X of the second convolution 2 ∈R C×H×W perform global average pooling to capture the overall distribution characteristics of the regions related to icing in the icing image, compress the spatial dimension of each channel C into a scalar, and keep the number of channels as C, but the spatial dimension is reduced to 1×1, that is:

[0033]

[0034] where, z c represents the global feature value of the c-th channel, and X 2 (i,j) are the pixel values at the i,j positions on the feature map X 2 and H×W is the size of the feature map;

[0035] Input the global feature vector Z∈R C into two fully connected layers. The first fully connected layer reduces the dimension, and then passes through the ReLU activation function:

[0036] Z′ = ReLU(W 1 Z)

[0037] Raise the dimension of the reduced feature back to the original number of channels C, and obtain the attention weight of each channel through the Sigmoid function:

[0038] S = σ(W 2 Z')

[0039] In the formula, W 1 and W 2 are the weight matrices of the first and second fully connected layers respectively, used to map the weights to between (0,1);

[0040] The weight S∈R C calculated by multiplying with the corresponding feature channels is reapplied to each channel of the feature map X 2 to adjust the weight size of each channel:

[0041] X′ 2 = X 2 ×S

[0042] Add the output feature map X′ 2 after channel attention weighting and the input feature map X element-wise, and apply a ReLU activation function to the output after addition to obtain the output of the residual block:

[0043] Y = ReLU(X′2 +X)

[0044] In the formula, X' 2 is the convolution feature map after adding channel attention.

[0045] Optionally, the image processing process of the decoder is specifically as follows:

[0046] The decoder uses the bilinear interpolation method to calculate new pixel values by performing two linear interpolations on the pixel values in the horizontal and vertical directions, and expands the size of the feature map from H×W to 2H×2W:

[0047] I(x,y) = (x 2 - x)(y 2 - y)I(x 1 ,y 1 )+(x - x 1 )(y 2 - y)I(x 2 ,y 1 )+(x 2 - x)(y - y 1 )I(x 1 ,y 2 )+(x - x 1 )(y - y 1 )I(x 2 ,y 2 )

[0048] In the formula, I(x 1 ,y 1 ), I(x 2 ,y 1 ), I(x 1 ,y 2 ), I(x 2 ,y 2 ) are the pixel points of the original feature map, I(x,y) is the value of the inserted pixel point, the inserted pixel point is at the position (x,y) and satisfies x 1 ≤x≤x 2 and y 1 ≤y≤y 2 ;

[0049] Perform a deconvolution operation on the interpolated feature map:

[0050]

[0051] Where \(x\) is the input feature map, \(K\) is the convolution kernel, \(Y(i,j)\) is the output feature map, \(s\) is the stride, \(i,j\) are the pixel positions of the output feature map, \(m,n\) are the indices of the convolution kernel. The spatial resolution of the input feature map is extended by inserting zero values, and the convolution kernel is used to perform convolution operations on the extended region to generate new pixel values, further enhancing the detailed information of the image;

[0052] The alternating operations of bilinear interpolation and transposed convolution are performed multiple times. A \(1\times1\) convolution kernel is used in the last layer of the decoder to generate an RGB image with 3 channels. The Tanh activation function is applied to limit the output pixel values within the range of \([-1,1]\) until the original resolution of the image is restored; The output layer formula is:

[0053]

[0054] Where is the dehazed image, \(x'\) is the feature map obtained from the alternating operations, and \(b\) is the bias term;

[0055] Skip connections are introduced after the multi-layer transposed convolution operations in the decoder:

[0056] F l concat = Concat(E l , D L-l )

[0057] Where \(F l concat is the concatenated feature map, \(E l is the output feature map of the \(l\)-th layer of the encoder, \(D L-l is the output feature map of the \((L - l)\)-th layer of the decoder, Concat represents the concatenation operation in the channel dimension, and the skip connection forms a new feature map by concatenating \(E l and \(D L-l in the channel dimension;

[0058] After concatenation, a convolution operation is used to fuse these feature maps:

[0059] F l out = Conv(F l concat )

[0060] Where \(F l out is the fused feature map.

[0061] Optionally, the true / false discrimination process of the discriminator is specifically as follows:

[0062] Input the generated image after dehazing by the generator and the real image \(I dehazed, the discriminator learns from the real image I dehazed and the generated image after defogging by the generator to enhance its ability to judge true and false images;

[0063] The discriminator uses the first-layer convolution to extract preliminary low-level features, and performs convolution on the input image using a 4×4 convolution kernel with a stride of 2:

[0064] y 1 = LeakyReLU(Conv2D(I, K 1 ) + b 1 )

[0065] Subsequent convolution layers each perform convolution using a 4×4 convolution kernel with a stride of 2. Through each layer of convolution, the spatial size of the feature map gradually shrinks:

[0066] y n+1 = LeakyReLU(Conv2D(y n , K n+1 ) + b n+1 )

[0067] In the formula, I is the input image of the discriminator, K 1 is the first-layer convolution kernel, b 1 is the bias, LeakyReLU is the activation function, K n+1 is the convolution kernel of the (n + 1)th layer, b n+1 is the bias of the (n + 1)th layer, and y n is the output feature map of the nth layer;

[0068] The last layer of the discriminator outputs a two-dimensional matrix:

[0069] D output = σ(Conv2D(y n , K n+1 ) + b n+1 )

[0070] In the formula, σ is the Sigmoid activation function. Each value of the two-dimensional matrix corresponds to a small block area in the input image for true / false judgment. After the input image I undergoes multiple convolution operations, the size of the output two-dimensional matrix is M×N, where each element represents the true / false prediction value of a k×k Patch.

[0071] Optionally, the loss function of the improved CycleGAN model generator is specifically:

[0072]

[0073] In the formula, They are the GAN adversarial loss, cycle consistency loss, feature consistency loss, and channel attention mechanism loss, respectively, λ cyc , λ id , λ att are respectively hyperparameters of, G is the generator, D X is the discriminator, I haze is the input foggy image, I real is the real fog-free image, E represents the loss expectation, F represents the inverse generator, ‖·‖ 1 represents the L1 norm, α c represents the attention weight of channel c, G c (I haze ) represents the output of the c-th channel in the generated image, I real,c is the true value of the c-th channel in the real image.

[0074] As can be seen from the above technical solutions, compared with the prior art, the present invention provides a method for removing fog from ice-covered transmission line images based on improved CycleGAN, which has the following beneficial effects:

[0075] Based on the CycleGAN structure, the present invention directly learns the mapping relationship from image data. Through adversarial training, the generator can generate high-quality fog-removed images. The present invention improves CycleGAN by introducing a channel attention mechanism. The channel attention mechanism can automatically focus on important feature channels in the image and weight and enhance these channels, so that key ice-covered information such as ice layer edges, transmission line surface details, ice thickness texture features, and transparency can be better retained in the fog-removed image, improving the clarity of the fog-removed image. The present invention is not only applicable to removing fog from ice-covered transmission line images under specific meteorological conditions, but also can effectively solve the problem of blurred ice-covered transmission line images under complex weather conditions, and can obtain fog-free images with higher clarity. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0077] Figure 1 is the flowchart of the method for removing fog from ice-covered transmission line images of the present invention;

[0078] Figure 2 is the schematic structural diagram of the improved CycleGAN model of the present invention;

[0079] Figure 3 This is a schematic diagram of the structure of the improved CycleGAN model generator of the present invention;

[0080] Figure 4 Schematic diagram of the residual block structure of the improved CycleGAN model of the present invention;

[0081] Figure 5 Schematic diagram of the channel attention mechanism structure introduced by the residual block of the present invention. DETAILED DESCRIPTION

[0082] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0083] The embodiment of the present invention discloses a method for defogging ice-covered images of power transmission lines based on an improved CycleGAN. Figure 1 As shown, the following steps are included:

[0084] S1, obtaining fog-free images and foggy images of ice-covered transmission lines, and preprocessing the images, including image denoising, image normalization, and image cropping adjustment;

[0085] S2. Build a generator of the improved CycleGAN model, including an encoder, a residual block, and a decoder connected in sequence, and input the preprocessed foggy image into the generator for learning;

[0086] S3. Construct a discriminator for the improved CycleGAN model. The discriminator determines the authenticity of the images generated by the generator and the real images.

[0087] S4. Use adversarial loss, cycle consistency loss, feature consistency loss and channel attention mechanism loss to construct the loss function of the improved CycleGAN model generator;

[0088] S5. Use foggy images of ice-covered power transmission lines under severe weather conditions and fog-free images under normal weather conditions to train the generator and discriminator. The generator gradually generates clearer defogging images through adversarial training and cycle consistency loss, and the discriminator judges the authenticity of the defogging images through adversarial training.

[0089] S6. After the training of the improved CycleGAN model is completed, the generator generates a defogged image based on the input foggy image.

[0090] In an embodiment of the present invention, the improved CycleGAN model includes two generators and two discriminators. As Figure 2 shown, generator G1 converts an image in the input domain X into an image in the target domain Y, and generator G2 performs the opposite task, converting an image in the target domain Y back into an image in the input domain X; discriminator D1 is used to discriminate whether an image comes from the target domain Y or is converted from the input domain X through G1, and discriminator D2 is used to discriminate whether an image comes from the input domain X or is converted from the target domain Y through G2. Both generator G1 and generator G2 include an encoder, a residual block, and a decoder connected in sequence.

[0091] Further, S1 is specifically as follows:

[0092] S11. Through an image acquisition device deployed on the transmission line, collect fog-free images of the ice-covered transmission line under normal climate conditions and foggy images under harsh meteorological conditions such as haze, snowfall, and low-temperature frost. The collected images include the transmission line itself, the surrounding environment, and the part covered by ice.

[0093] The collected ice-covered images should include the transmission line itself, the surrounding environment, and the part covered by ice. The imaging device needs to have a high resolution to capture sufficient details in scenes with large light contrasts. To ensure the usability of the input images and improve the defogging effect of the ice-covered images, after collecting the ice-covered images, it is necessary to preprocess the images, including image denoising, image normalization, image cropping, and adjustment.

[0094] S12. Since the images may be affected by the external environment (such as dust, raindrops, wind, and snow) during the collection of ice-covered images, the images will contain noise. Denoise the collected images through Gaussian filtering. Gaussian filtering smooths the images to reduce the noise introduced by environmental and equipment factors, making the overall details of the images more coherent.

[0095] In an embodiment of the present invention, the Gaussian filtering formula is as follows:

[0096]

[0097] In the formula, (x, y) is the position coordinate relative to the central pixel within the filtering window, and σ is the standard deviation of the Gaussian function.

[0098] S13. Due to the change of environmental illumination, there are large differences in the brightness and contrast of the ice-covered images. Therefore, perform pixel normalization processing on the collected images, divide each pixel value of the images by 255, and scale the pixel values of the images to the [0, 1] interval to ensure that the subsequent defogging model can better process images with different brightnesses.

[0099] S14. To ensure that the defogging model can process input images of a fixed size, the icing images are all cropped to a size of 256×256 pixels, and the images are rotated or flipped according to actual needs to highlight the icing part of the transmission line. Crop the images according to the image content to remove redundant background information. To ensure that the defogging model can process input images of a fixed size, the images are all cropped to a size of 256×256 pixels to ensure that the transmission line and its icing area are in prominent positions in the image;

[0100] In the embodiment of the present invention, in S14, the images are also rotated or flipped according to actual needs to highlight the icing part of the transmission line.

[0101] Further, as Figure 3 shown, the image generation process of the generator is specifically as follows:

[0102] The encoder uses multiple convolutional layers to extract low-level features of the input image. In the convolutional layer, the input image is downsampled through convolutional operations, gradually reducing the size of the feature map and extracting deep-level features. The convolutional layer contains multiple convolutional kernels, and the convolutional kernels slide on the input feature map based on dot product operations to calculate the feature values at each position, thereby extracting specific spatial features;

[0103] Use residual blocks to extract low-level and high-level features in the input image through convolutional layers. Calculate the weights of each channel through global average pooling and fully connected layers of the channel attention mechanism, select the channels containing information on the key icing areas, enhance their features, and perform residual addition on the weighted feature map and the original feature map to form a residual connection;

[0104] In the decoder, bilinear interpolation is used as the initial upsampling operation, and the feature map is further refined through deconvolution operations, enhancing the detail retention effect of defogging while ensuring the smoothness of upsampling.

[0105] Bilinear interpolation can quickly and smoothly expand the size of the feature map. Since bilinear interpolation can maintain the overall smoothness of the image and deconvolution can capture more detailed features through the learning of convolutional kernels, the method of combining bilinear interpolation and deconvolution is used for image restoration, enhancing the detail retention effect of defogging while ensuring the smoothness of upsampling.

[0106] Further, the image processing process of the encoder is specifically as follows:

[0107] The first layer of the encoder uses a convolutional layer to extract the initial features of the image, performs batch normalization on the output after convolution to accelerate convergence and improve training stability, and uses the ReLU activation function to introduce non-linearity and enhance the feature expression ability;

[0108] In the embodiments of the present invention, the received input image is usually an RGB image with a size of H×W×3, where H is the height of the image, W is the width of the image, and 3 represents the RGB channels; the first-layer convolutional kernel is 7×7, the stride is set to 1, and the padding is 3, so that the size of the output feature map remains unchanged, that is:

[0109] F 1 =Conv(l, k = 7, s = 1, p = 3)

[0110] In the formula, l is the input image, k is the convolutional kernel size, s is the stride, p is the padding, and Conv represents the convolution operation; the batch normalization and ReLU activation function are specifically:

[0111] F 1 ' = ReLU(BN(F 1 ))

[0112] ReLU(x) = max(0, x)

[0113] Among them, BN represents batch normalization.

[0114] The second layer of the encoder uses a convolutional layer for downsampling to further extract features, so that the size of the output feature map is reduced by half. Batch normalization is performed on the output after convolution, and the ReLU activation function is used;

[0115] In the embodiments of the present invention, the second-layer convolutional kernel is 3×3, the stride is set to 2, and the padding is 1, so that the size of the output feature map is reduced by half, that is:

[0116] F 2 =Conv(F 1 ', k = 3, s = 2, p = 1)

[0117] The batch normalization and ReLU activation function are specifically:

[0118] F 2 ' = ReLU(BN(F 2 ))

[0119] The third layer of the encoder uses a convolutional layer to further downsample the feature map, extract higher-level features, and further reduce the resolution of the feature map. Batch normalization is performed on the output after convolution and the ReLU activation function is used.

[0120] In the embodiments of the present invention, the third-layer convolutional kernel is a 3×3 convolutional layer with a stride of 2, which further reduces the resolution of the feature map, that is:

[0121] F 3 =Conv(F 2 ', k = 3, s = 2, p = 1)

[0122] The batch normalization and ReLU activation function are specifically as follows:

[0123] F 3 ′ = ReLU(BN(F 2 ))

[0124] Through the convolution operations of 3 convolutional layers, the resolution of the feature map can be gradually reduced, and more abundant feature information can be extracted.

[0125] Furthermore, as Figure 4 shown, the image processing process of the residual block is specifically as follows:

[0126] Convolve the input feature map X, and then perform batch normalization and ReLU activation:

[0127] X 1 = ReLU(BN(Conv(X)))

[0128] Perform another convolution and batch normalization operation without using the activation function:

[0129] X 2 = BN(Conv(X 1 ))

[0130] In the formula, X 1 is the feature map output by the first convolution, X 2 is the feature map output by the second convolution, BN represents batch normalization, and Conv represents convolution;

[0131] In the embodiment of the present invention, the convolution kernel of the convolutional layer of the residual block is 3×3, and the stride is 1.

[0132] As Figure 5 shown, perform global average pooling on the output X 2 ∈R C×H×W of the second convolution to capture the overall distribution characteristics of the regions related to icing in the icing image, compress the spatial dimension of each channel C into a scalar, and keep the number of channels as C, but the spatial dimension is reduced to 1×1, that is:

[0133]

[0134] Among them, z c represents the global feature value of the c-th channel, X 2 (i,j) is the pixel value at the i,j position on the feature map X 2 , and H×W is the size of the feature map;

[0135] The global feature vector Z∈R CInput to two fully connected layers. The first fully connected layer reduces the dimension, and then passes through the ReLU activation function:

[0136] Z′ = ReLU(W 1 Z)

[0137] In the embodiment of the present invention, the dimension reduction factor is 16;

[0138] The dimension-reduced features are re-expanded to the original number of channels C, and the attention weight of each channel is obtained through the Sigmoid function:

[0139] S = σ(W 2 Z')

[0140] Wherein, W 1 and W 2 are the weight matrices of the first and second fully connected layers respectively, used to map the weights to between (0,1), indicating the importance of the channels;

[0141] The weight S ∈ R calculated by multiplying with the corresponding feature channels C , and the learned channel weight S is reapplied to each channel of the feature map X 2 to adjust the weight size of each channel:

[0142] X′ 2 = X 2 × S

[0143] So that the channels related to the ice layer edge, ice cover thickness, morphology and position will obtain higher weights, while the irrelevant or noise information will be suppressed;

[0144] Finally, a residual connection is performed to retain the preliminarily extracted low-level features and strengthen the important features related to ice coverage. The output feature map X′ 2 after channel attention weighting and the input feature map X are added element by element, and a ReLU activation function is applied to the output after addition to obtain the output of the residual block:

[0145] Y = ReLU(X′ 2 + X)

[0146] Wherein, X′ 2 is the convolutional feature map after adding channel attention.

[0147] The residual block structure disclosed in the embodiment of the present invention can not only improve the image quality after defogging, but also better retain the key details in the transmission line ice-covered image, and improve the accurate recognition and processing of the ice-covered situation.

[0148] Furthermore, the image processing process of the decoder is specifically as follows:

[0149] The decoder uses the method of bilinear interpolation to calculate the new pixel values by performing two linear interpolations on the pixel values in the horizontal and vertical directions, expanding the size of the feature map from H×W to 2H×2W:

[0150] I(x,y) = (x 2 -x)(y 2 -y)I(x 1 ,y 1 )+(x-x 1 )(y 2 -y)I(x 2 ,y 1 )+(x 2 -x)(y-y 1 )I(x 1 ,y 2 )+(x-x 1 )(y-y 1 )I(x 2 ,y 2 )

[0151] In the formula, I(x 1 ,y 1 ), I(x 2 ,y 1 ), I(x 1 ,y 2 ), I(x 2 ,y 2 ) are the pixel points of the original feature map, I(x,y) is the value of the inserted pixel point, the inserted pixel point is at the position (x,y) and satisfies x 1 ≤x≤x 2 and y 1 ≤y≤y 2 ;

[0152] The above interpolation formula obtains the value of the target pixel point by weighted summation of four adjacent pixel points. The weight is determined by the distance between the target point and the four pixel points. For each insertion position, the interpolation pixel value is calculated according to the formula to generate a high-resolution image;

[0153] Perform a deconvolution operation on the interpolated feature map:

[0154]

[0155] In the formula, x is the input feature map, K is the convolution kernel, Y(i,j) is the output feature map, s is the stride, i,j are the pixel positions of the output feature map, m,n are the indices of the convolution kernel. The spatial resolution of the input feature map is expanded by inserting zero values, and the convolution kernel is used to perform a convolution operation on the expanded region to generate new pixel values, further improving the detail information of the image;

[0156] The alternating operations of bilinear interpolation and deconvolution are performed multiple times. A 1×1 convolutional kernel is used in the last layer of the decoder to generate an RGB image with 3 channels. The Tanh activation function is applied to limit the output pixel values within the range of [-1, 1] until the original resolution of the image is restored. The output layer formula is:

[0157]

[0158] where, is the dehazed image, x′ is the feature map obtained from the alternating operations, and b is the bias term;

[0159] Since the low-level features in the encoder may contain useful icing edge information, texture features, and local details, and these information may be partially lost in the layer-by-layer deconvolution, therefore, skip connections are introduced after the multi-layer deconvolution operations in the decoder:

[0160] F l concat = Concat(E l , D L-l )

[0161] where, F l concat is the concatenated feature map, E l is the output feature map of the l-th layer of the encoder, D L-l is the output feature map of the (L - l)-th layer of the decoder, Concat represents the concatenation operation in the channel dimension, and the skip connection forms a new feature map by concatenating E l and D L-l in the channel dimension; by directly passing the low-level features to the high-level of the decoder, the skip connection retains these detail information and can generate images with higher quality;

[0162] After concatenation, a convolutional operation is used to fuse these feature maps:

[0163] F l out = Conv(F l concat )

[0164] where, F l out is the fused feature map. Through multi-layer skip connections, the decoder can generate dehazed images with rich details and finally output clear and high-resolution dehazed images.

[0165] In the dehazing model based on CycleGAN, after the generator generates a high-resolution haze-free image, the discriminator is used to distinguish between the "dehazed image" output by the generator and the "real haze-free image". Through adversarial training, the generator continuously optimizes its output to generate more realistic images, while the discriminator continuously improves its ability to distinguish between real images and generated images. The discriminator part uses PatchGAN in CycleGAN as the discriminator, which divides the input image into multiple small regions and judges the authenticity of each region, enabling the generator to generate images with richer details;

[0166] Furthermore, the process of the discriminator's true / false discrimination is specifically as follows:

[0167] Input the generated image after dehazing by the generator and the real image I dehazed , the discriminator learns from the real image I dehazed and the generated image after dehazing by the generator to enhance its ability to judge true and false images;

[0168] The discriminator uses the first layer of convolution to extract preliminary low-level features and performs convolution on the input image using a 4×4 convolutional kernel with a stride of 2:

[0169] y 1 = LeakyReLU(Conv2D(I, K 1 ) + b 1 )

[0170] Subsequent convolutional layers each use a 4×4 convolutional kernel with a stride of 2 for convolution. Through each layer of convolution, the spatial size of the feature map gradually shrinks:

[0171] y n+1 = LeakyReLU(Conv2D(y n , K n+1 ) + b n+1 )

[0172] In the formula, I is the input image of the discriminator, K 1 is the first layer of convolutional kernel, b 1 is the bias, LeakyReLU is the activation function, K n+1 is the (n + 1)-th layer of convolutional kernel, b n+1 is the (n + 1)-th layer of bias, and y n is the output feature map of the n-th layer;

[0173] The last layer of the discriminator outputs a two-dimensional matrix:

[0174] D output = σ(Conv2D(y n , K n+1) + b n+1 )

[0175] In the formula, σ is the Sigmoid activation function. Each value of the two-dimensional matrix corresponds to a small patch area in the input image for true or false judgment. After the input image I undergoes multiple convolutional operations, the output two-dimensional matrix has a size of M×N, where each element represents the true or false prediction value of a k×k Patch.

[0176] Furthermore, the loss function of the improved CycleGAN model generator is specifically:

[0177]

[0178] In the formula, are the GAN adversarial loss, cycle consistency loss, feature consistency loss, and channel attention mechanism loss respectively. λ cyc 、λ id 、λ att are the hyperparameters of respectively. G is the generator, D X is the discriminator, I haze is the input foggy image, I real is the real fog-free image, E represents the loss expectation, F represents the inverse generator, ‖·‖ 1 represents the L1 norm, α c represents the attention weight of channel c, G c (I haze ) represents the output of the c-th channel in the generated image, and I real,c is the true value of the c-th channel in the real image.

[0179] The loss function directly affects the optimization direction of the generator and the discriminator, ensuring that the generator can generate realistic fog-free images while maintaining the feature consistency between the input and output images. After introducing the channel attention mechanism, the design of the loss function needs to enhance the model's attention to key features, that is, the contribution of important feature channels can be more effectively highlighted through the attention mechanism. Therefore, the generator loss function of the dehazing model is composed of a combination of adversarial loss, cycle consistency loss, feature consistency loss, and channel attention mechanism. Through the combination of these loss functions, the model can generate high-quality dehazing images while maintaining the consistency of input and output features;

[0180] The adversarial loss function is used to optimize the generator so that the generated dehazing images can deceive the discriminator. The adversarial loss of CycleGAN adopts the GAN adversarial loss, and through the game training of the generator G and the discriminator D, the generator is optimized to generate real fog-free images;

[0181] The cyclic consistency loss function is used to ensure that the generator does not overly change the structure and content of the input image. In the defogging task, the generator not only needs to generate a fog-free image, but also needs to ensure the consistency of the generated image with the input image in terms of structure, features, etc.;

[0182] The feature consistency loss is used to constrain the output image of the generator to be consistent with the input image in appearance, avoiding image distortion caused by excessive defogging. In the defogging model, this loss can help the generator retain the details of the input image;

[0183] After introducing the channel attention mechanism, the loss function needs to enhance the model's attention to the key feature channels. The channel attention mechanism can automatically learn the weights of each channel and perform weighting according to importance.

[0184] Furthermore, S5 is trained using transmission line icing images collected under various adverse weather conditions (such as haze, snowfall, etc.). During the training process, the generator and the discriminator are alternately optimized. The generator continuously adjusts the defogging image generation process to make the generated image approximate the real defogged image; the discriminator, by distinguishing the difference between the generated defogged image and the real fog-free image, prompts the generator to continuously improve.

[0185] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0186] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for defogging ice-covered images of power transmission lines based on improved CycleGAN, characterized in that: The following steps are involved: S1, obtaining fog-free images and foggy images of ice-covered transmission lines, and preprocessing the images, including image denoising, image normalization, and image cropping adjustment; S2. Build a generator of the improved CycleGAN model, including an encoder, a residual block, and a decoder connected in sequence, and input the preprocessed foggy image into the generator for learning; S3. Construct a discriminator for the improved CycleGAN model. The discriminator determines the authenticity of the images generated by the generator and the real images. S4. Use adversarial loss, cycle consistency loss, feature consistency loss and channel attention mechanism loss to construct the loss function of the improved CycleGAN model generator; S5. Use foggy images of ice-covered power transmission lines under severe weather conditions and fog-free images under normal weather conditions to train the generator and discriminator. The generator gradually generates clearer defogging images through adversarial training and cycle consistency loss, and the discriminator judges the authenticity of the defogging images through adversarial training. S6. After the training of the improved CycleGAN model is completed, the generator generates a defogged image based on the input foggy image.

2. According to claim 1, a method for defogging ice-covered images of power transmission lines based on improved CycleGAN is characterized in that: S1 is specifically: S11. Using image acquisition equipment deployed on the transmission line, collect fog-free images of the ice-covered transmission line under normal weather conditions and foggy images under severe weather conditions, wherein the collected images include the transmission line itself, the surrounding environment, and the part covered by ice; S12, denoising the collected image by using Gaussian filtering. Gaussian filtering reduces the noise introduced by environmental and equipment factors by smoothing the image, making the overall details of the image more coherent; S13, performing pixel normalization processing on the collected image, dividing each pixel value of the image by 255, so that the pixel value of the image is scaled to the interval [0,1], to ensure that the later defogging model can better process images of different brightness; S14, cropping the image according to the image content, removing redundant background information, and cropping the image to a size of 256×256 pixels.

3. According to claim 1, a method for defogging ice-covered images of power transmission lines based on improved CycleGAN is characterized in that: The image generation process of the generator is as follows: The encoder uses multiple convolutional layers to extract low-level features of the input image. In the convolutional layer, the input image is downsampled through convolution operations, the feature map size is gradually reduced, and deep features are extracted. The convolutional layer contains multiple convolution kernels, which slide on the input feature map based on the dot product operation to calculate the feature value of each position, thereby extracting specific spatial features. The residual block is used to extract low-level and high-level features from the input image through the convolution layer. The weight of each channel is calculated through the global average pooling and fully connected layer of the channel attention mechanism. The channel containing the information of the key ice-covered area is selected and its features are enhanced. The residual of the weighted feature map is added to the original feature map to form a residual connection. In the decoder, bilinear interpolation is used as a preliminary upsampling operation, and the feature map is further refined through deconvolution operations, which ensures the smoothness of upsampling while enhancing the detail retention effect of dehazing.

4. According to claim 3, the method for defogging ice-covered images of power transmission lines based on improved CycleGAN is characterized in that: The image processing process of the encoder is as follows: The first layer of the encoder uses a convolutional layer to extract the initial features of the image, performs batch normalization on the output after convolution, and uses the ReLU activation function to introduce nonlinearity to enhance the feature expression capability; The second layer of the encoder uses a convolutional layer for downsampling to further extract features, so that the size of the output feature map is reduced by half, the output after convolution is batch normalized, and the ReLU activation function is used; The third layer of the encoder uses a convolutional layer to further downsample the feature map, extract higher-level features, further reduce the resolution of the feature map, batch normalize the convolutional output and use the ReLU activation function.

5. According to claim 3, the method for defogging ice-covered images of power transmission lines based on improved CycleGAN is characterized in that: The image processing process of the residual block is specifically as follows: Convolution is performed on the input feature map X, followed by batch normalization and ReLU activation: X1 = ReLU(BN(Conv(X))) Perform another convolution and batch normalization operation without using an activation function: X2=BN(Conv(X1)) In the formula, X1 is the feature map output by the first convolution, X2 is the feature map output by the second convolution, BN means batch normalization, and Conv means convolution; The output of the second convolution X2∈R C×H×W Perform global average pooling to capture the overall distribution characteristics of the ice-related areas in the ice-covered image, compress the spatial dimension of each channel C into a scalar, keep the number of channels as C, but reduce the spatial dimension to 1×1, that is: Among them, z c represents the global eigenvalue of the cth channel, X2(i,j) is the pixel value at position i,j on the feature map X2, and H×W is the feature map size; The global eigenvector Z∈R C The input is sent to two fully connected layers, the first fully connected layer reduces the dimension, and then passes through the ReLU activation function: Z′=ReLU(W1Z) The reduced-dimensional features are re-dimensionalized back to the original number of channels C, and the attention weight of each channel is obtained through the Sigmoid function: S=σ(W2Z') Where W1 and W2 are the weight matrices of the first and second fully connected layers, respectively, used to map the weights to between (0,1); The weight S∈R is calculated by multiplying the corresponding feature channel C , reapply the learned channel weight S to each channel of the feature map X2, and adjust the weight size of each channel: X′2=X2×S The output feature map X′2 weighted by channel attention and the input feature map X are added element by element, and a ReLU activation function is applied to the added output to obtain the output of the residual block: Y = ReLU(X′2+X) Where X′2 is the convolutional feature map after adding channel attention.

6. According to claim 3, the method for defogging ice-covered images of power transmission lines based on improved CycleGAN is characterized in that: The image processing process of the decoder is as follows: The decoder uses bilinear interpolation to calculate new pixel values ​​by performing two linear interpolations on the pixel values ​​in the horizontal and vertical directions, expanding the size of the feature map from H×W to 2H×2W: I(x,y)=(x2-x)(y2-y)I(x1,y1)+(x-x1)(y2-y)I(x2,y1)+(x2-x)(y-y1)I(x1,y2)+(x-x1)(y-y1)I(x2,y2) Where I(x1,y1), I(x2,y1)I(x1,y2), I(x2,y2) are the pixels of the original feature map, I(x,y) is the value of the inserted pixel, the inserted pixel is at position (x,y) and satisfies x1≤x≤x2 and y1≤y≤y2; Perform deconvolution on the interpolated feature map: In the formula, x is the input feature map, K is the convolution kernel, Y(i,j) is the output feature map, s is the step size, i,j are the pixel positions of the output feature map, m,n are the indexes of the convolution kernel, and the spatial resolution of the input feature map is expanded by inserting zero values. The convolution kernel is used to perform convolution operations on the expanded area to generate new pixel values, thereby further improving the detail information of the image. Perform alternating bilinear interpolation and deconvolution operations multiple times. In the last layer of the decoder, a 1×1 convolution kernel is used to generate an RGB image with 3 channels. The Tanh activation function is applied to limit the output pixel value to the range of [-1, 1] until the original resolution of the image is restored. The output layer formula is: In the formula, is the dehazed image, x′ is the feature map obtained by alternating operations, and b is the bias term; Introduce skip connections after the multi-layer deconvolution operation of the decoder: F l concat =Concat(E l ,D L-l ) In the formula, F l concat is the concatenated feature map, E l is the output feature map of the encoder layer l, D L-l is the output feature map of the Llth layer of the decoder, Concat represents the concatenation operation on the channel dimension, and the skip connection is achieved by concatenating E l and D L-l Splice to form a new feature map; After concatenation, a convolution operation is used to fuse these feature maps: F l out =Conv(F l concat ) In the formula, F l out is the fused feature map.

7. According to claim 1, a method for defogging ice-covered images of power transmission lines based on improved CycleGAN is characterized in that: The true and false discrimination process of the discriminator is specifically as follows: Input the generated image after dehazing and the real image I dehazed , the discriminator is trained from the real image I dehazed And the generated image after the generator is dehazed Learning from the image, enhancing its ability to distinguish true from false images; The discriminator uses the first layer of convolution to extract preliminary low-level features and convolves the input image with a 4×4 convolution kernel with a stride of 2: y1=LeakyReLU(Conv2D(I,K1)+b1) Each subsequent convolutional layer uses a 4×4 convolution kernel with a stride of 2 for convolution. Through each layer of convolution, the spatial size of the feature map is gradually reduced: y n+1 =LeakyReLU(Conv2D(y n ,K n+1 )+b n+1 ) Where I is the input image of the discriminator, K1 is the first layer convolution kernel, b1 is the bias, LeakyReLU is the activation function, and K n+1 is the convolution kernel of the n+1th layer, b n+1 is the bias of the n+1th layer, y n is the output feature map of the nth layer; The output of the last layer of the discriminator is a two-dimensional matrix: D output =σ(Conv2D(y n ,K n+1 )+b n+1 ) Where σ is the Sigmoid activation function. Each value of the two-dimensional matrix corresponds to a small area in the input image for true or false judgment. After the input image undergoes multi-layer convolution operations, the output two-dimensional matrix size is M×N, where each element represents the true or false prediction value of a k×k Patch.

8. According to claim 1, the method for defogging ice-covered images of power transmission lines based on improved CycleGAN is characterized in that: The loss function of the improved CycleGAN model generator is specifically: In the formula, They are GAN adversarial loss, cycle consistency loss, feature consistency loss, channel attention mechanism loss, λ cyc , id , att They are Hyperparameters, G is the generator, D X is the discriminator, I haze is the input foggy image, I real is a real haze-free image, E represents the loss expectation, F represents the inverse generator, ‖·‖1 represents the L1 norm, and α c represents the attention weight of channel c, G c (I haze ) represents the output of the cth channel in the generated image, I real,c is the true value of the cth channel in the real image.