A method and system for denoising transmission line inspection images
By introducing Mix module, CSM module and CCAM module into the SADNet denoising network model, the SADNet-S denoising network model is constructed, which solves the detection difficulty caused by complex environments and morphological changes in insulator detection, and achieves the improvement of efficient image denoising and target detection accuracy.
Patent Information
- Application Number
- CN202510073511.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The prior art faces complex environments, many background interferences, many changes in the shape and color of the insulator, large changes in scale, rotation and posture, and detection difficulty under different lighting conditions, resulting in low detection accuracy.
An improvement solution based on SADNet denoising network model is proposed, and the Mix module, CSM module and CCAM module are introduced to build a SADNet-S denoising network model. Through multi-scale feature extraction, dynamic adaptive filtering and attention mechanism, the image denoising performance and object detection accuracy are improved.
Effectively remove image noise in complex environments, retain key features, significantly improve insulator detection accuracy, can cope with insulator morphology and color changes in different environments, and improve detection robustness and generalization.
Smart Images

Figure CN119540566B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power transmission lines, and in particular relates to a method and system for denoising a power transmission line inspection image. Background Art
[0002] Insulators are a key component of the power system, responsible for supporting and isolating high-voltage wires. Their stability is directly related to the safety and reliability of power transmission. With the increase in power demand and the increasing complexity of power systems, the timely detection and accurate location of insulator faults have become particularly important.
[0003] Traditional insulator detection methods rely on manual inspections and regular checks, but these methods face problems of low efficiency and interference from human factors, and it is difficult to conduct timely and effective detection at the early stage of a fault. Although manual inspection can find some obvious defects, it cannot meet the needs of real-time monitoring and efficient screening. In order to solve this problem, image processing methods based on computer vision and deep learning have gradually become a research hotspot in the field of insulator detection.
[0004] Although current image denoising technology has great potential in insulator detection, it still faces some challenges. First, the transmission line environment is complex, with a lot of background interference, and the shape and color of insulators vary greatly, which increases the difficulty of detection. In addition, the scale of insulators in images varies greatly, and existing target detection algorithms may have unstable performance when dealing with multi-scale targets. Finally, the rotation, posture changes of insulators in images, and the performance under different lighting conditions may also affect the detection effect. Summary of the invention
[0005] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a method and system for denoising images of transmission line inspections, and to propose an efficient image denoising solution, which aims to improve the quality of insulator images and the accuracy of subsequent target detection, and provide support for intelligent inspections of power systems.
[0006] To achieve the above object, the present invention provides the following technical solution: a method for denoising a transmission line inspection image, comprising the following steps:
[0007] S1: Get the original image of the insulator, add noise to the original image, and build an insulator noise image dataset;
[0008] S2: Based on the SADNet denoising network model, the Mix module, CSM module and CCAM module are introduced to build the SADNet-S denoising network model;
[0009] The processing process of the SADNet-S denoising network model is as follows: in the encoding stage, the input image is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain the feature map conv1, the feature map conv1 is down-sampled by the stride convolution block to obtain the feature map pool1, the feature map pool1 is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain the feature map conv2, the feature map conv2 is down-sampled by the stride convolution block to obtain the feature map pool2, the feature map pool2 is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain The feature map conv3 is downsampled by the stride convolution block to obtain the feature map pool3, and the feature map pool3 is subjected to residual convolution by the residual convolution block to obtain the feature map conv4. The feature map conv4 captures multi-scale information through the context extraction module and is further enhanced by the CSM module to obtain the enhanced feature map conv4'; in the decoding stage, the first offset block generates an offset and applies it to the enhanced feature map conv4', and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv4, and then through the multi-scale convolution operation and D The FSA module generates feature maps conv_y and conv_z of different scales, and uses the CCAM module to fuse the feature maps conv_y and conv_z to obtain the enhanced feature map dconv4'. The enhanced feature map dconv4' is upsampled and concatenated with the feature map conv3 to obtain the fused feature map up3. The second offset block generates an offset and applies it to the feature map conv3, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv3. The feature map dconv3 is upsampled and concatenated with the feature map conv2 to obtain the fused feature map up2. The third offset block generates an offset and applies it to the feature map conv2, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv2. The feature map dconv2 is upsampled and concatenated with the feature map conv1 to obtain the fused feature map up1. The fourth offset block generates an offset and applies it to the feature map conv1, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv1. Finally, the convolutional layer is used to compress the number of channels to the output number of channels, and it is added pixel by pixel to the input image to obtain the denoised output image.
[0010] S3: Use the insulator noise image dataset to train the SADNet-S denoising network model and obtain the optimal training weights;
[0011] S4: Load the optimal training weights into the SADNet-S denoising network model, and input the insulator noise image to be tested into the SADNet-S denoising network model to obtain the denoised insulator image.
[0012] Further preferably, noise is added to the original image, and the pixel value I of the image after adding noise is noisy The formula for (x,y) is:
[0013] ,
[0014] Where I(x,y) is the pixel value of the original image at position (x,y), x and y represent the horizontal and vertical coordinates of the pixel in the original image, respectively, and N(x,y) is the Gaussian noise of the same size as the original image;
[0015] The generated Gaussian noise N(x,y) follows the following distribution:
[0016] ,
[0017] The mean is 0 and the standard deviation is Gaussian distribution.
[0018] Further preferably, the processing process of the Mix module is as follows: the input features are normalized by batch normalization, and then features are extracted by point convolution and convolution operations, and then the features are divided into multiple sub-blocks in the channel dimension according to the preset number of groups, each sub-block is subjected to independent convolution and batch normalization, and the outputs of multiple sub-blocks are re-joined in the channel dimension to obtain spliced features, and the spliced features are subjected to point convolution-RELU-point convolution operations, combined with Sigmoid function to extract features and then residually connected with the spliced features, and Softmax is used to normalize between blocks to obtain weight distributions of different blocks, and the features are point-by-point multiplied with the corresponding sub-block features and then spliced to obtain fused features, and the fused features are processed in parallel by three different sizes of dilated convolutions to obtain three parallel features, and the three parallel features are spliced and processed by a multilayer perceptron MLP, which includes two point convolutions and a GELU activation function, and the output of the multilayer perceptron MLP is residually connected with the input features and then output to the parallel attention module EPA;
[0019] The parallel attention module EPA includes a simple pixel attention module, a channel attention module and a pixel attention module. The processing process of the parallel attention module EPA is as follows: the input features of the parallel attention module EPA are normalized by batch normalization, and the normalized features are processed in parallel by the simple pixel attention module, the channel attention module and the pixel attention module; the simple pixel attention module contains two branches: one is the feature extraction branch PFs, and the other is the pixel gating branch PAs. The PFs branch extracts features by point convolution and convolution operations, while the PAs branch uses point convolution and Sigmoid activation function to calculate the importance of each pixel. Finally, the outputs of the PFs branch and the PAs branch are element-wise multiplied to obtain the final feature F s; the channel attention module extracts the global channel gating features through global average pooling, point convolution-GELU-point convolution operations, combined with the Sigmoid function, and the global channel gating features are element-wise multiplied with the input features of the channel attention module to obtain the final feature Gs; the pixel attention module extracts the global pixel gating features through point convolution-GELU-point convolution operations, combined with the Sigmoid function, and the global pixel gating features are element-wise multiplied with the input features of the pixel attention module to obtain the final feature Hs; the features obtained by the simple pixel attention module, the pixel attention module and the channel attention module are merged, and then MLP processing is performed through the multi-layer perceptron, and finally residual connection is performed with the input features of the parallel attention module EPA to obtain the output.
[0020] Further preferably, the CSM module includes a CABSE module, a SABSE module and an MMSDC module connected in sequence; the CABSE module includes two parallel branches, both of which are connected to average pooling and maximum pooling respectively, and then input into the SE attention module, and output the channel attention weight after the GELU activation function and 1×1 convolution processing, and finally add the results of the two parallel branches and use the Sigmoid function to obtain the final weight map, and the final weight map is multiplied element by element with the input of the CABSE module to obtain the output; the SABSE module includes two parallel branches, both of which are passed through the channel maximum pool The output of the two parallel branches is processed by the Sigmoid function to generate a spatial attention map, and then multiplied element-by-element with the input of the SABSE module to obtain the output; the input of the MMSDC module is split into several paths, each path is processed by the SE attention module and batch normalization, and then split into two branches, one branch is processed by the ReLU6 activation function, and the other branch is processed by the deep convolution and GELU activation function. The outputs of all branches are merged at the channel or feature level, and the information of different branches is further fused and scattered through channel shuffling to obtain the output.
[0021] Further preferably, the CCAM module adopts the U-shaped structure of the encoder and decoder of CasDyF-Net, including a DFSA module, a LLFB module and a RMB module. The processing process of the CCAM module is: first, the DLK module performs channel alignment and preliminary processing on the input features to obtain a unified feature representation; then, the features pass through multiple DFSA modules in sequence, each DFSA module splits the channel into several sub-blocks according to a preset ratio, and performs convolution, batch normalization and attention operations on each sub-block respectively to obtain the weight distribution corresponding to each sub-block, and multiplies the weight corresponding to each sub-block point by point to the sub-block feature, then all sub-blocks are re-spliced in the channel dimension to obtain a fusion feature that fuses multi-scale context information, and then the fusion feature enters the LLFB module, first performs global average pooling to extract the global statistical information of each channel, and then performs local interaction and fusion between the sub-blocks through convolution and attention mechanisms. After completing local fusion, the global information is integrated through the global receptive field operation, and the channel attention is used to integrate the global information. The attention mechanism focuses on key channels, captures the feature expression of salient areas through the spatial attention mechanism, and then optimizes and reconstructs the global features in combination with 1×1 convolution blocks. Finally, the original input features are added to the attention-enhanced features through the RMB module. The processing process of the RMB module is as follows: the input features of the RMB module are first processed through three parallel branches. The first parallel branch uses 3x3 dilated convolution and PReLU activation function, and is residually connected to the input features of the RMB module. The second parallel branch uses 3x3 standard convolution and PReLU activation function, and is residually connected to the output of the first parallel branch. The third parallel branch uses 3x3 dilated convolution and PReLU activation function, and is residually connected to the output of the second parallel branch. Then, the input features of the RMB module and the outputs of the three parallel branches are spliced in the channel dimension, and the spliced features are fused through two layers of point convolution and GELU activation function, and finally are residually connected to the input features of the RMB module to obtain the output.
[0022] Further preferably, the loss function of the SADNet-S denoising network model is expressed as:
[0023] ,
[0024] In the formula, is the loss function of the SADNet-S denoising network model, For the The actual value of the observation point, For the The model prediction value of observation points, N is the total number of observation points, is the actual value, is the model prediction value, It is the first layer feature map, is the noise variance, It is a global coefficient used to adjust the weight of the entire loss function. It is an important hyperparameter for adjusting perceptual loss and pixel loss. It is used to balance the regularization term , n is the number of observation points for which errors need to be calculated, and L is the number of layers in the perceptual network used to calculate the perceptual loss.
[0025] A transmission line inspection image denoising system, comprising a data set construction module, a model construction module, a model training module, and an output module;
[0026] Dataset construction module: obtain the original image of the insulator, add noise to the original image, and construct the insulator noise image dataset;
[0027] Model building module: Based on the SADNet denoising network model, the Mix module, CSM module and CCAM module are introduced to build the SADNet-S denoising network model;
[0028] The processing process of the SADNet-S denoising network model is as follows: in the encoding stage, the input image is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain the feature map conv1, the feature map conv1 is down-sampled by the stride convolution block to obtain the feature map pool1, the feature map pool1 is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain the feature map conv2, the feature map conv2 is down-sampled by the stride convolution block to obtain the feature map pool2, the feature map pool2 is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain The feature map conv3 is downsampled by the stride convolution block to obtain the feature map pool3, and the feature map pool3 is subjected to residual convolution by the residual convolution block to obtain the feature map conv4. The feature map conv4 captures multi-scale information through the context extraction module and is further enhanced by the CSM module to obtain the enhanced feature map conv4'; in the decoding stage, the first offset block generates an offset and applies it to the enhanced feature map conv4', and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv4, and then through the multi-scale convolution operation and D The FSA module generates feature maps conv_y and conv_z of different scales, and uses the CCAM module to fuse the feature maps conv_y and conv_z to obtain the enhanced feature map dconv4'. The enhanced feature map dconv4' is upsampled and concatenated with the feature map conv3 to obtain the fused feature map up3. The second offset block generates an offset and applies it to the feature map conv3, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv3. The feature map dconv3 is upsampled and concatenated with the feature map conv2 to obtain the fused feature map up2. The third offset block generates an offset and applies it to the feature map conv2, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv2. The feature map dconv2 is upsampled and concatenated with the feature map conv1 to obtain the fused feature map up1. The fourth offset block generates an offset and applies it to the feature map conv1, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv1. Finally, the convolutional layer is used to compress the number of channels to the output number of channels, and it is added pixel by pixel to the input image to obtain the denoised output image.
[0029] Model training module: Use the insulator noise image dataset to train the SADNet-S denoising network model to obtain the optimal training weights;
[0030] Output module: Load the optimal training weights into the SADNet-S denoising network model, and input the insulator noise image to be tested into the SADNet-S denoising network model to obtain the denoised insulator image.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] The present invention proposes a SADNet-S denoising network model based on multi-scale feature extraction, dynamic adaptive filtering and attention mechanism improvement. By adopting the SADNet-S denoising network model, its powerful generalization and robustness are fully utilized to achieve efficient image denoising and significantly improve the target detection accuracy. The SADNet-S denoising network model can effectively remove image noise and retain key features in complex environments through its high-precision denoising performance, thereby providing a clearer input image for target detection. This greatly improves the detection accuracy of insulators under different shooting angles and lighting conditions. In addition, the high robustness of the SADNet-S denoising network model enables it to effectively cope with changes in the morphology and color of insulators in different environments, and can maintain a high denoising accuracy regardless of background interference or scale changes of objects, especially in complex scenes, such as tasks such as power transmission line inspection. The SADNet-S denoising network model can combine effective denoising with target detection, and can more accurately capture the true shape and position of objects, especially when dealing with objects with rotation and posture changes, which can significantly improve the detection accuracy. Through this innovative network design, the SADNet-S denoising network model not only performs excellently in accuracy, but also has strong robustness and generalization. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flow chart of the method of the present invention;
[0034] Figure 2 This is a schematic diagram of the SADNet-S denoising network model structure;
[0035] Figure 3 It is a schematic diagram of the Mix module structure;
[0036] Figure 4 This is a schematic diagram of the SPA structure of a simple pixel attention module;
[0037] Figure 5 This is a schematic diagram of the channel attention module CA structure;
[0038] Figure 6 Schematic diagram of the pixel attention module PA structure;
[0039] Figure 7 It is a schematic diagram of the CSM module structure;
[0040] Figure 8 It is a schematic diagram of the CABSE module structure;
[0041] Fig. 9 It is a schematic diagram of the SABSE module structure;
[0042] Fig.10 It is a schematic diagram of the MMSDC module structure;
[0043] Fig.11 It is a schematic diagram of the CCAM module structure;
[0044] Fig.12 It is a schematic diagram of the DFSA module structure;
[0045] Fig.13 This is a schematic diagram of the LLFB module structure. DETAILED DESCRIPTION
[0046] The present invention is further described below in conjunction with embodiments. It is necessary to point out that the following embodiments are only used to further illustrate the present invention and cannot be understood as limiting the scope of protection of the present invention. Some non-essential improvements and adjustments made by technical personnel in this field based on the above invention content still fall within the scope of protection of the present invention.
[0047] Example 1: Please refer to Figure 1 As shown, a method for denoising a transmission line inspection image in this embodiment includes the following steps:
[0048] S1: Get the original image of the insulator, add noise to the original image, and build an insulator noise image dataset;
[0049] Add noise to the original image. The pixel value of the image after adding noise is I noisy The formula for (x,y) is:
[0050] ,
[0051] Where I(x,y) is the pixel value of the original image at position (x,y) (usually the value of the three RGB channels), x and y represent the horizontal and vertical coordinates of the pixel in the original image, respectively, and N(x,y) is the Gaussian noise of the same size as the original image;
[0052] The generated Gaussian noise N(x,y) follows the following distribution:
[0053] ,
[0054] The mean is 0 and the standard deviation is Gaussian distribution.
[0055] S2: Based on the SADNet denoising network model, the SADNet-S denoising network model is constructed; the Mix module is introduced to address the problem that the original SADNet denoising network model relies on small convolution kernels and standard convolution operations, which leads to limited expression capabilities when processing complex textures and structures. The CSM module is introduced to address the problem that the original SADNet denoising network model lacks a detailed attention mechanism and cannot fully utilize the spatial and channel information of the image, resulting in unsatisfactory denoising effects. The CCAM module is introduced to address the problem that the original SADNet denoising network model cannot fully utilize the inter-channel dependencies due to the lack of a channel attention mechanism, resulting in difficulty in capturing complex or subtle features.
[0056] like Figure 2As shown in Figure 1, the processing process of the SADNet-S denoising network model is as follows: in the encoding stage, the input image is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain the feature map conv1, the feature map conv1 is down-sampled by the stride convolution block to obtain the feature map pool1, the feature map pool1 is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain the feature map conv2, the feature map conv2 is down-sampled by the stride convolution block to obtain the feature map pool2, the feature map pool2 is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain the feature map conv3, the feature map conv3 is down-sampled by the stride convolution block to obtain the feature map pool3, and the feature map pool ol3 performs residual convolution through the residual convolution block to obtain the feature map conv4. The feature map conv4 captures multi-scale information through the context extraction module and is further enhanced by the CSM module to obtain the enhanced feature map conv4'. In the decoding stage, the first offset block generates an offset and applies it to the enhanced feature map conv4' to achieve deformable convolution. Then the enhanced feature map conv4' is formed after the RSABlock residual attention mechanism extracts features, and then the feature maps conv_y and conv_z of different scales are generated through multi-scale convolution operations and the DFSA module. The feature map conv_y and the feature map conv_z are fused using the CCAM module to obtain the enhanced feature. Figure dconv4', upsample the enhanced feature map dconv4' and concatenate it with the feature map conv3 to obtain the fused feature map up3, the second offset block generates an offset and applies it to the feature map conv3, then extracts features through the RSABlock residual attention mechanism to form the feature map dconv3, upsamples the feature map dconv3 and concatenates it with the feature map conv2 to obtain the fused feature map up2, the third offset block generates an offset and applies it to the feature map conv2, then extracts features through the RSABlock residual attention mechanism to form the feature map dconv2, upsamples the feature map dconv2 and concatenates it with the feature map conv1 to obtain the fused feature map up1, the fourth offset block generates an offset The offset is applied to the feature map conv1, and then the feature map dconv1 is formed after the RSABlock residual attention mechanism extracts the features. Finally, the convolution layer is used to compress the number of channels to the output channel number, and the channel number is added to the input image pixel by pixel to obtain the denoised output image. The processing process of the deformable convolution module is as follows: first, it is determined whether the number of input channels is equal to the number of output channels. If they are not equal, the number of channels is adjusted by point convolution. Then, the input features and offset features are extracted by deformable convolution DCN. The deformable convolution uses a 3x3 convolution kernel, a step size of 1, a padding of 1, a void rate of 1, and a deformable group number of 8. The output of the deformable convolution is activated by the LeakyReLU activation function (negative slope is 0.2) Perform nonlinear transformation; then extract features through 3x3 convolution (step size 1, padding 1), and finally perform residual connection with the input features to obtain the output. The convolution layer weights in the module are initialized using Xavier uniform distribution, and the bias term is initialized to 0. The processing process of the transposed convolution module is as follows: the input features are first upsampled through the first transposed convolution to reduce the number of input channels from 8 times the baseline number of channels to 4 times the baseline number of channels, and a 2x2 convolution kernel and a step size of 2 are used to upsample the features, so that the spatial size of the output feature map is expanded to twice that of the input feature map; then the second transposed convolution is continued to upsample, reducing the number of channels from 4 times the baseline number of channels to 2 times the baseline number of channels, and the 2x2 convolution kernel and step size of 2 are used to further expand the feature map size by 2 times; finally, the third transposed convolution is used for the final upsampling, reducing the number of channels from 2 times the baseline number of channels to the baseline number of channels, and the 2x2 convolution kernel and step size of 2 are still used to expand the feature map spatial size by 2 times again. The processing process of the offset block is as follows: the input feature is first extracted through the first convolutional layer offset_conv1 to obtain the initial offset feature. If there is offset information last_offset of the previous layer, it is upsampled to 2 times the size by bilinear interpolation, and the upsampled offset information is multiplied by 2 and then concatenated with the current initial offset feature in the channel dimension. The concatenated feature continues to extract features through the second convolutional layer offset_conv2; if there is no offset information of the previous layer, the initial offset feature is directly input into the third convolutional layer offset_conv3 for processing; the output of all convolutional layers is nonlinearly transformed through the LeakyReLU activation function, where the slope of the negative semi-axis is 0.2; all convolutional layer weights in the offset block are initialized using Xavier uniform distribution, and the bias term is initialized to zero; finally, the processed offset feature is output to guide subsequent feature alignment operations. .
[0057] like Figure 3As shown in the figure, the processing process of the Mix module is as follows: the input features are normalized by batch normalization, and then the features are extracted by point convolution and convolution operations. Then, according to the preset number of groups, the features are divided into multiple sub-blocks in the channel dimension. Each sub-block undergoes independent convolution and batch normalization. The outputs of multiple sub-blocks are re-joined in the channel dimension to obtain the spliced features. The spliced features are extracted by point convolution-RELU-point convolution operations, combined with the Sigmoid function, and then residually connected with the spliced features. Softmax is used between the blocks. Normalize to get the weight distribution of different blocks, multiply them point by point with the corresponding sub-block features and then concatenate them to get the fused features. The fused features are processed in parallel by three different sizes of dilated convolutions (5x5, 3x3, and 7x7 convolution kernels, and the dilation rate is 3) to get three parallel features. The three parallel features are concatenated and processed by a multi-layer perceptron MLP. The multi-layer perceptron MLP contains two point convolutions and a GELU activation function. The output of the multi-layer perceptron MLP is residually connected with the input features and then output to the parallel attention module EPA.
[0058] The parallel attention module EPA includes a simple pixel attention module SPA, a channel attention module CA and a pixel attention module PA. The processing process of the parallel attention module EPA is as follows: the input features of the parallel attention module EPA are normalized by batch normalization, and the normalized features are processed in parallel by the simple pixel attention module SPA, the channel attention module CA and the pixel attention module PA; Figure 4 As shown in the figure, the simple pixel attention module SPA consists of two branches: one is the feature extraction branch PFs, and the other is the pixel gated branch PAs. The PFs branch extracts features through point convolution and convolution operations, while the PAs branch uses point convolution and Sigmoid activation function to calculate the importance of each pixel. Finally, the outputs of the PFs branch and the PAs branch are element-wise multiplied to obtain the final feature Fs; as shown in the figure Figure 5 As shown in , the channel attention module CA extracts the global channel gating feature through global average pooling, point convolution-GELU-point convolution operation, combined with the Sigmoid function. The global channel gating feature is element-wise multiplied with the input feature of the channel attention module to obtain the final feature Gs; Figure 6As shown in the figure, the pixel attention module PA extracts the global pixel gating feature through the point convolution-GELU-point convolution operation combined with the Sigmoid function. The global pixel gating feature is element-wise multiplied with the input feature of the pixel attention module to obtain the final feature Hs; the features obtained by the simple pixel attention module SPA, the pixel attention module PA and the channel attention module CA are merged, and then processed by the multi-layer perceptron MLP, and finally residually connected with the input feature of the parallel attention module EPA to obtain the output.
[0059] like Figure 7 As shown, the CSM module includes a CABSE module, a SABSE module and an MMSDC module connected in sequence; Figure 8 As shown in , the CABSE module includes two parallel branches, both of which are connected to average pooling and maximum pooling respectively, and then input into the SE attention module, and output the channel attention weight after the GELU activation function and 1×1 convolution processing, and finally add the results of the two parallel branches and use the Sigmoid function to obtain the final weight map, and the final weight map is multiplied element by element with the input of the CABSE module to obtain the output; as shown in Fig. 9 As shown in Figure 1, the SABSE module includes two parallel branches. Both parallel branches are processed by the channel maximum pooling and SE attention module, and then processed by a variety of different convolution combinations. The outputs of the two parallel branches are generated by the Sigmoid function to generate a spatial attention map, which is then multiplied element-by-element with the input of the SABSE module to obtain the output; Fig.10 As shown in the figure, the input of the MMSDC module is divided into several paths. Each path passes through the SE attention module and batch normalization, and then is split into two branches. One branch is processed by the ReLU6 activation function, and the other branch is processed by deep convolution and GELU activation function. The outputs of all branches are merged at the channel or feature level, and the output is obtained by further fusion and dispersion of the information of different branches through channel shuffling.
[0060] The CCAM module adopts the U-shaped structure of the encoder and decoder of CasDyF-Net, such as Fig.11 As shown, including DFSA module (such as Fig.12 As shown), LLFB module (such as Fig.13As shown in the figure, and the RMB module, the processing process of the CCAM module is as follows: first, the DLK module performs channel alignment and preliminary processing on the input features to obtain a unified feature representation; then, the features pass through multiple DFSA modules in sequence, each DFSA module splits the channel into several sub-blocks according to a preset ratio, and performs convolution, batch normalization and attention operations (including Softmax normalization) on each sub-block respectively, to obtain the weight distribution corresponding to each sub-block, and multiply the weight corresponding to each sub-block point by point to the sub-block feature to achieve differentiated weighting; then all sub-blocks are re-joined in the channel dimension to obtain a fused feature that fuses multi-scale contextual information, and then the fused feature enters the LLFB module, first performs global average pooling to extract the global statistical information of each channel, and then performs local interaction and fusion between the sub-blocks through the convolution and attention mechanism, which not only retains the richness of multi-scale features, but also makes the information between sub-blocks more tightly coupled. After completing local fusion, the global information is integrated through the global receptive field operation, and the channel attention mechanism is used to focus on the key channels. The feature expression of the salient area is captured through the spatial attention mechanism, and the global features are optimized and reconstructed in combination with the 1×1 convolution block. Finally, the original input features are added to the features after attention enhancement through the RMB module. The processing process of the RMB module is as follows: the input features of the RMB module are first processed through three parallel branches. The first parallel branch uses 3x3 dilated convolution (with a dilated rate of 5) and PReLU activation function, and is combined with the input features of the RMB module. Residual connection; the second parallel branch uses 3x3 standard convolution (padding is 1) and PReLU activation function, and is residually connected to the output of the first parallel branch; the third parallel branch uses 3x3 void convolution (void ratio is 3) and PReLU activation function, and is residually connected to the output of the second parallel branch; then the input features of the RMB module and the outputs of the three parallel branches are concatenated in the channel dimension, and the concatenated features are fused through two layers of point convolution and GELU activation function, and finally residually connected with the input features of the RMB module to obtain the output.
[0061] The loss function of the SADNet-S denoising network model is expressed as:
[0062] ,
[0063] In the formula, is the loss function of the SADNet-S denoising network model, For the The actual value of the observation point, For the The model prediction value of observation points, N is the total number of observation points, is the actual value, is the model prediction value, It is the first layer feature map, is the noise variance, It is a global coefficient used to adjust the weight of the entire loss function. It is an important hyperparameter for adjusting perceptual loss and pixel loss. It is used to balance the regularization term , n is the number of observation points for which errors need to be calculated, and L is the number of layers in the perceptual network used to calculate the perceptual loss.
[0064] S3: Use the insulator noise image dataset to train the SADNet-S denoising network model and obtain the optimal training weights;
[0065] The insulator noise image dataset in S1 is used as training data, and the SADNet-S denoising network model is iteratively trained 100 times using the default hyperparameter settings. After the training is completed, the optimal training weights of the SADNet-S denoising network model are obtained.
[0066] S4: Load the optimal training weights into the SADNet-S denoising network model, and input the insulator noise image to be tested into the SADNet-S denoising network model to obtain the denoised insulator image.
[0067] In summary, the present invention proposes a SADNet-S denoising network model based on multi-scale feature extraction, dynamic adaptive filtering and attention mechanism improvement. By adopting the SADNet-S denoising network model, its powerful generalization and robustness are fully utilized to achieve efficient image denoising and significantly improve the target detection accuracy. The SADNet-S denoising network model can effectively remove image noise and retain key features in complex environments through its high-precision denoising performance, thereby providing a clearer input image for target detection. This greatly improves the detection accuracy of insulators under different shooting angles and lighting conditions. In addition, the high robustness of the SADNet-S denoising network model enables it to effectively cope with changes in the morphology and color of insulators in different environments, whether it is background interference or scale changes of objects, it can maintain a high denoising accuracy, especially in complex scenes, such as tasks such as power line inspection. The SADNet-S denoising network model can combine effective denoising with target detection, and can more accurately capture the true shape and position of objects, especially when dealing with objects with rotation and posture changes, which can significantly improve the detection accuracy. Through this innovative network design, the SADNet-S denoising network model not only performs excellently in accuracy, but also has strong robustness and generalization.
[0068] Embodiment 2: A transmission line inspection image denoising system described in this embodiment includes a data set construction module, a model construction module, a model training module, and an output module;
[0069] Dataset construction module: obtain the original image of the insulator, add noise to the original image, and construct the insulator noise image dataset;
[0070] Model building module: Based on the SADNet denoising network model, the Mix module, CSM module and CCAM module are introduced to build the SADNet-S denoising network model;
[0071] The processing process of the SADNet-S denoising network model is as follows: in the encoding stage, the input image is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain the feature map conv1, the feature map conv1 is down-sampled by the stride convolution block to obtain the feature map pool1, the feature map pool1 is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain the feature map conv2, the feature map conv2 is down-sampled by the stride convolution block to obtain the feature map pool2, the feature map pool2 is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain The feature map conv3 is downsampled by the stride convolution block to obtain the feature map pool3, and the feature map pool3 is subjected to residual convolution by the residual convolution block to obtain the feature map conv4. The feature map conv4 captures multi-scale information through the context extraction module and is further enhanced by the CSM module to obtain the enhanced feature map conv4'; in the decoding stage, the first offset block generates an offset and applies it to the enhanced feature map conv4', and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv4, and then through the multi-scale convolution operation and D The FSA module generates feature maps conv_y and conv_z of different scales, and uses the CCAM module to fuse the feature maps conv_y and conv_z to obtain the enhanced feature map dconv4'. The enhanced feature map dconv4' is upsampled and concatenated with the feature map conv3 to obtain the fused feature map up3. The second offset block generates an offset and applies it to the feature map conv3, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv3. The feature map dconv3 is upsampled and concatenated with the feature map conv2 to obtain the fused feature map up2. The third offset block generates an offset and applies it to the feature map conv2, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv2. The feature map dconv2 is upsampled and concatenated with the feature map conv1 to obtain the fused feature map up1. The fourth offset block generates an offset and applies it to the feature map conv1, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv1. Finally, the convolutional layer is used to compress the number of channels to the output number of channels, and it is added pixel by pixel to the input image to obtain the denoised output image.
[0072] Model training module: Use the insulator noise image dataset to train the SADNet-S denoising network model to obtain the optimal training weights;
[0073] Output module: Load the optimal training weights into the SADNet-S denoising network model, and input the insulator noise image to be tested into the SADNet-S denoising network model to obtain the denoised insulator image.
[0074] The above only expresses the preferred implementation of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosure to modify or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.
Claims
1. A method for denoising a transmission line inspection image, characterized in that: The following steps are involved: S1: Get the original image of the insulator, add noise to the original image, and build an insulator noise image dataset; S2: Based on the SADNet denoising network model, the Mix module, CSM module and CCAM module are introduced to build the SADNet-S denoising network model; The processing process of the SADNet-S denoising network model is as follows: in the encoding stage, the input image is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain the feature map conv1, the feature map conv1 is down-sampled by the stride convolution block to obtain the feature map pool1, the feature map pool1 is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain the feature map conv2, the feature map conv2 is down-sampled by the stride convolution block to obtain the feature map pool2, the feature map pool2 is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain The feature map conv3 is downsampled by the stride convolution block to obtain the feature map pool3, and the feature map pool3 is subjected to residual convolution by the residual convolution block to obtain the feature map conv4. The feature map conv4 captures multi-scale information through the context extraction module and is further enhanced by the CSM module to obtain the enhanced feature map conv4'; in the decoding stage, the first offset block generates an offset and applies it to the enhanced feature map conv4', and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv4, and then through the multi-scale convolution operation and D The FSA module generates feature maps conv_y and conv_z of different scales, and uses the CCAM module to fuse the feature maps conv_y and conv_z to obtain the enhanced feature map dconv4'. The enhanced feature map dconv4' is upsampled and concatenated with the feature map conv3 to obtain the fused feature map up3. The second offset block generates an offset and applies it to the feature map conv3, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv3. The feature map dconv3 is upsampled and concatenated with the feature map conv2 to obtain the fused feature map up2. The third offset block generates an offset and applies it to the feature map conv2, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv2. The feature map dconv2 is upsampled and concatenated with the feature map conv1 to obtain the fused feature map up1. The fourth offset block generates an offset and applies it to the feature map conv1, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv1. Finally, the convolutional layer is used to compress the number of channels to the output number of channels, and it is added pixel by pixel to the input image to obtain the denoised output image. S3: Use the insulator noise image dataset to train the SADNet-S denoising network model and obtain the optimal training weights; S4: Load the optimal training weights into the SADNet-S denoising network model, and input the noise image of the insulator to be tested into the SADNet-S denoising network model to obtain the denoised insulator image; The processing process of the Mix module is as follows: the input features are normalized by batch normalization, and then features are extracted by point convolution and convolution operations. Then, according to the preset number of groups, the features are divided into multiple sub-blocks in the channel dimension, each sub-block is subjected to independent convolution and batch normalization processing, and the outputs of multiple sub-blocks are re-joined in the channel dimension to obtain spliced features. The spliced features are extracted by point convolution-RELU-point convolution operations combined with Sigmoid functions and then residually connected with the spliced features. Softmax is used to normalize between blocks to obtain weight distributions of different blocks, and the features are multiplied point by point with the corresponding sub-block features and then spliced to obtain fused features. The fused features are processed in parallel by three different sizes of dilated convolutions to obtain three parallel features. The three parallel features are spliced and processed by a multi-layer perceptron MLP. The multi-layer perceptron MLP includes two point convolutions and a GELU activation function. The output of the multi-layer perceptron MLP is residually connected with the input features and then output to the parallel attention module EPA. The parallel attention module EPA includes a simple pixel attention module, a channel attention module and a pixel attention module. The processing process of the parallel attention module EPA is as follows: the input features of the parallel attention module EPA are normalized by batch normalization, and the normalized features are processed in parallel by the simple pixel attention module, the channel attention module and the pixel attention module; the simple pixel attention module contains two branches: one is the feature extraction branch PFs, and the other is the pixel gating branch PAs. The PFs branch extracts features by point convolution and convolution operations, while the PAs branch uses point convolution and Sigmoid activation function to calculate the importance of each pixel. Finally, the outputs of the PFs branch and the PAs branch are element-wise multiplied to obtain the final feature F s; the channel attention module extracts the global channel gating feature through global average pooling, point convolution-GELU-point convolution operation, combined with the Sigmoid function, and the global channel gating feature is element-wise multiplied with the input feature of the channel attention module to obtain the final feature Gs; the pixel attention module extracts the global pixel gating feature through point convolution-GELU-point convolution operation, combined with the Sigmoid function, and the global pixel gating feature is element-wise multiplied with the input feature of the pixel attention module to obtain the final feature Hs; the features obtained by the simple pixel attention module, the pixel attention module and the channel attention module are merged, and then MLP is performed through the multi-layer perceptron, and finally the residual connection is performed with the input feature of the parallel attention module EPA to obtain the output; The CSM module includes a CABSE module, a SABSE module and an MMSDC module connected in sequence; the CABSE module includes two parallel branches, both of which are connected to average pooling and maximum pooling respectively, and then input into the SE attention module, and output the channel attention weight after the GELU activation function and 1×1 convolution processing, and finally add the results of the two parallel branches and use the Sigmoid function to obtain the final weight map, and the final weight map is multiplied element by element with the input of the CABSE module to obtain the output; the SABSE module includes two parallel branches, both of which are subjected to channel maximum pooling and S The output of the two parallel branches is processed by the SABSE module and then processed by a variety of different convolution combinations. The output of the two parallel branches generates a spatial attention map through the Sigmoid function and is multiplied element-by-element with the input of the SABSE module to obtain the output. The input of the MMSDC module is split into several paths, each of which is processed by the SE attention module and batch normalization, and then split into two branches. One branch is processed by the ReLU6 activation function, and the other branch is processed by the deep convolution and GELU activation function. The outputs of all branches are merged at the channel or feature level, and the information of different branches is further fused and scattered through channel shuffling to obtain the output. The CCAM module adopts the U-shaped structure of the encoder and decoder of CasDyF-Net, including a DFSA module, a LLFB module and an RMB module. The processing process of the CCAM module is as follows: first, the DLK module performs channel alignment and preliminary processing on the input features to obtain a unified feature representation; then, the features pass through multiple DFSA modules in sequence, each DFSA module splits the channel into several sub-blocks according to a preset ratio, and performs convolution, batch normalization and attention operations on each sub-block respectively to obtain the weight distribution corresponding to each sub-block, and multiplies the weight corresponding to each sub-block point by point to the sub-block feature, then all sub-blocks are re-spliced in the channel dimension to obtain a fusion feature that fuses multi-scale context information, and then the fusion feature enters the LLFB module, first performs global average pooling to extract the global statistical information of each channel, and then performs local interaction and fusion between each sub-block through convolution and attention mechanism. After completing local fusion, the global information is integrated through the global receptive field operation, and the channel attention mechanism is used to integrate the global information. The key channels are focused, the feature expression of the salient areas is captured through the spatial attention mechanism, and the global features are optimized and reconstructed in combination with the 1×1 convolution block. Finally, the original input features are added to the features after attention enhancement through the RMB module. The processing process of the RMB module is as follows: the input features of the RMB module are first processed through three parallel branches. The first parallel branch uses 3x3 dilated convolution and PReLU activation function, and is residually connected with the input features of the RMB module. The second parallel branch uses 3x3 standard convolution and PReLU activation function, and is residually connected with the output of the first parallel branch. The third parallel branch uses 3x3 dilated convolution and PReLU activation function, and is residually connected with the output of the second parallel branch. Then, the input features of the RMB module and the outputs of the three parallel branches are spliced in the channel dimension, and the spliced features are feature fused through two layers of point convolution and GELU activation function, and finally are residually connected with the input features of the RMB module to obtain the output.
2. The power transmission line inspection image denoising method according to claim 1, characterized in that: Add noise to the original image. The pixel value of the image after adding noise is I noisy The formula for (x,y) is: , Where I(x,y) is the pixel value of the original image at position (x,y), x and y represent the horizontal and vertical coordinates of the pixel in the original image, respectively, and N(x,y) is the Gaussian noise of the same size as the original image; The generated Gaussian noise N(x,y) follows the following distribution: , The mean is 0 and the standard deviation is Gaussian distribution.
3. The method for denoising a power transmission line inspection image according to claim 1, characterized in that: The expression of the loss function of the SADNet-S denoising network model is: , In the formula, is the loss function of the SADNet-S denoising network model, For the The actual value of the observation point, For the The model prediction value of observation points, N is the total number of observation points, is the actual value, is the model prediction value, It is the first layer feature map, is the noise variance, It is a global coefficient used to adjust the weight of the entire loss function. It is an important hyperparameter for adjusting perceptual loss and pixel loss. It is used to balance the regularization term , n is the number of observation points for which errors need to be calculated, and L is the number of layers in the perceptual network used to calculate the perceptual loss.
4. A transmission line inspection image denoising system, used to implement the denoising method according to any one of claims 1 to 3, characterized in that: Including data set construction module, model construction module, model training module, and output module; Dataset construction module: obtain the original image of the insulator, add noise to the original image, and construct the insulator noise image dataset; Model building module: Based on the SADNet denoising network model, the Mix module, CSM module and CCAM module are introduced to build the SADNet-S denoising network model; The processing process of the SADNet-S denoising network model is as follows: in the encoding stage, the input image is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain the feature map conv1, the feature map conv1 is down-sampled by the stride convolution block to obtain the feature map pool1, the feature map pool1 is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain the feature map conv2, the feature map conv2 is down-sampled by the stride convolution block to obtain the feature map pool2, the feature map pool2 is multi-scaled and attention enhanced by the residual convolution block and the Mix module to obtain The feature map conv3 is downsampled by the stride convolution block to obtain the feature map pool3, and the feature map pool3 is subjected to residual convolution by the residual convolution block to obtain the feature map conv4. The feature map conv4 captures multi-scale information through the context extraction module and is further enhanced by the CSM module to obtain the enhanced feature map conv4'; in the decoding stage, the first offset block generates an offset and applies it to the enhanced feature map conv4', and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv4, and then through the multi-scale convolution operation and D The FSA module generates feature maps conv_y and conv_z of different scales, and uses the CCAM module to fuse the feature maps conv_y and conv_z to obtain the enhanced feature map dconv4'. The enhanced feature map dconv4' is upsampled and concatenated with the feature map conv3 to obtain the fused feature map up3. The second offset block generates an offset and applies it to the feature map conv3, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv3. The feature map dconv3 is upsampled and concatenated with the feature map conv2 to obtain the fused feature map up2. The third offset block generates an offset and applies it to the feature map conv2, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv2. The feature map dconv2 is upsampled and concatenated with the feature map conv1 to obtain the fused feature map up1. The fourth offset block generates an offset and applies it to the feature map conv1, and then extracts features through the RSABlock residual attention mechanism to form the feature map dconv1. Finally, the convolutional layer is used to compress the number of channels to the output number of channels, and it is added pixel by pixel to the input image to obtain the denoised output image. Model training module: Use the insulator noise image dataset to train the SADNet-S denoising network model to obtain the optimal training weights; Output module: Load the optimal training weights into the SADNet-S denoising network model, and input the insulator noise image to be tested into the SADNet-S denoising network model to obtain the denoised insulator image.
Citation Information
Patent Citations
Moving image deblurring method based on self-adaptive residual errors and recursive cross attention
CN112164011A
Diffusion model defogging method fusing parallel multi-convolution attention
CN117994167A