A method for detecting pin defects in power transmission lines based on dynamic receptive field
By using dynamic receptive field and space activation module methods in the detection of pin defects of transmission line, the existing detection methods are solved, and higher detection accuracy and recognition rate are achieved.
Patent Information
- Application Number
- CN202210793757.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-07
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-07-07
AI Technical Summary
The existing transmission line pin defect detection methods have low accuracy and poor robustness. Especially in complex scenarios, changes in distance, height and low lead to changes in defect size, and increase network depth leads to loss of feature information.
Using a detection method based on dynamic receptive fields, a feature pyramid network combines receptive fields and channel information of different sizes, integrates context information, and builds embedded dynamic receptive fields modules and spatial activation modules to enhance the acquisition of information in the region of interest.
It improves the accuracy and recognition rate of pin defect detection in transmission line, enhances the robustness of the model, and can better adapt to defect detection in complex scenarios.
Smart Images

Figure CN115082798B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of power transmission line image analysis, and in particular to a method for detecting pin defects in a power transmission line based on a dynamic receptive field. Background Art
[0002] With the rapid development of drone aerial photography technology, transmission line images obtained through aerial photography can provide more comprehensive observation perspectives, efficiently capture the visual status of transmission lines while also having extremely high safety. According to the State Grid standards, defects in transmission line inspections mainly include tower defects, insulator defects, and hardware defects. In contrast, pins are small in size and numerous, and are widely found on towers and connecting hardware. Pins are generally installed at the connection points of various components of transmission lines to stabilize the structure. Due to the large force, problems such as pins falling off and pins loosening are prone to occur, which can easily cause large-scale power outages in the power system and affect power transmission safety.
[0003] However, the currently proposed pin defect detection methods have the following problems: (1) The pin defect scene of the transmission line is complex, and the changes in distance and height when shooting by drones will cause the size of the pin defect to change; (2) As the network depth increases, the defect features will lose a lot of information, and the region of interest generated by the network will also be affected.
[0004] Therefore, how to integrate contextual information and better extract pin defect features has become a technical problem that needs to be solved urgently. Summary of the invention
[0005] The purpose of the present invention is to solve the defects of low accuracy and poor robustness of the power transmission line pin defect detection method in the prior art, and to provide a power transmission line pin defect detection method based on dynamic receptive field to solve the above problems.
[0006] In order to achieve the above object, the technical solution of the present invention is as follows:
[0007] A method for detecting pin defects in a power transmission line based on a dynamic receptive field comprises the following steps:
[0008] Acquisition and preprocessing of transmission line pin defect samples: Acquire transmission line pin defect samples and perform data cleaning and data enhancement preprocessing on the defect samples;
[0009] Construction of a transmission line pin defect detection model: Construction of a transmission line pin defect detection model consisting of a feature extraction network embedded in a dynamic receptive field module to fuse context and inter-channel information, a region generation network embedded in a spatial activation module, and a defect detection network;
[0010] Training of the transmission line pin defect detection model: sending the preprocessed transmission line pin defect samples to the transmission line pin defect detection model for training;
[0011] Obtaining the image of the pins of the power transmission line to be detected: obtaining the image of the pins of the power transmission line to be detected and performing preprocessing;
[0012] Obtaining the transmission line pin defect detection results: input the preprocessed transmission line pin image to be detected into the trained transmission line pin defect detection model to obtain the transmission line pin defect detection results.
[0013] The construction of the transmission line pin defect detection model includes the following steps:
[0014] Based on the feature pyramid, a feature extraction network is constructed that combines dynamic receptive fields of different sizes and integrates context information;
[0015] Setting of feature pyramid: Setting the feature pyramid includes three parts: bottom-up, top-down and horizontal connection;
[0016] Set up a dynamic receptive field module: Set up a dynamic receptive field module consisting of two branches, namely, a large receptive field branch and a small receptive field branch with receptive fields of different sizes that fuse context and channel information;
[0017] By activating the spatial information of the region of interest, a spatially activated region generation network is constructed;
[0018] Setting of the region generation network: The region generation network is set to consist of two branches, a 1*1*18 convolution classification branch using softmax loss, and a 1*1*36 convolution regression branch using smooth-L1 loss. The region of interest is generated through the two branches;
[0019] The spatial activation module is set to consist of two branches, which are placed after each region of interest to further activate the spatial information of these regions. Specifically, one branch averages the pixel values of the corresponding points of the feature maps of all channels to obtain a feature map, and the global information of the feature layer is better obtained through averaging. The other branch takes the maximum value of the pixel values of the corresponding points of the feature maps of all channels to obtain another feature map, and the texture information of the feature layer is better obtained by taking the maximum value. Then the two feature maps are spliced together, and then activated by a 3*3 convolution layer and a sigmoid nonlinear function to generate a heat map and feed it back to the corresponding region of interest;
[0020] f(x)=σ{c[mean(x)+max(x)]}*x,
[0021] Among them, c is 3*3 convolution, σ is sigmoid activation function, x is the input region of interest, f(x) is the region of interest after activation, max(x) is the maximum value function, and mean(x) is the average value function;
[0022] Constructing a defect detection network: The defect detection network makes the final prediction of the region of interest after spatial activation, which includes two branches: a 1*1*C convolution classification branch using softmax loss and a 1*1*4 convolution positioning branch using smooth-L1 loss.
[0023] The training of the transmission line pin defect detection model includes the following steps:
[0024] The power transmission line pin defect detection model is set to be trained for 24 rounds in total, and the training data set image H*W is input into the feature extraction network, where H is the image height and W is the image width;
[0025] The training dataset image H*W enters the bottom-up part of the feature extraction network: the input image consists of 5 convolutional layers, each of which consists of 64-channel 1*1 convolution, 64-channel 3*3 convolution and 256-channel 1*1 convolution, generating four feature layers: C2 feature layer, C3 feature layer, C4 feature layer and C5 feature layer respectively;
[0026] Then enter the top-down and lateral connection parts of the feature extraction network:
[0027] Specifically, the C5 feature layer is first subjected to 1*1 convolution to obtain the M5 feature layer; M5 is first upsampled by 2 times, and then the C4 feature layer is connected horizontally by 1*1 convolution to obtain the M4 feature layer; the M4 feature layer is first upsampled by 2 times, and then the C3 feature layer is connected horizontally by 1*1 convolution to obtain the M3 feature layer; the M3 feature layer is first upsampled by 2 times, and then the C2 feature layer is connected horizontally by 1*1 convolution to obtain the M2 feature layer; the M2 feature layer, M3 feature layer, M4 feature layer, and M5 feature layer are then subjected to 3*3 convolution respectively to obtain the final P2 feature layer, P3 feature layer, P4 feature layer, and P5 feature layer;
[0028] The P2 feature layer, P3 feature layer, P4 feature layer, and P5 feature layer output by the feature extraction network are input into the region generation network of the embedded spatial activation module to generate regions of interest on the four feature layers of P2 feature layer, P3 feature layer, P4 feature layer, and P5 feature layer respectively;
[0029] The region of interest obtained by the region generation network embedded in the spatial activation module is cut out from the P2 feature layer, P3 feature layer, P4 feature layer, and P5 feature layer through a pooling layer and input into the defect detection network. The fully connected branch of the defect detection network uses cross entropy loss, and the bounding box regression branch uses GIOU loss as the loss function.
[0030] The setting of the feature pyramid includes the following steps:
[0031] The bottom-up part of the feature pyramid is set to pass through 5 convolutional layers. The 5 convolutional layers correspond to the convolutions with step sizes of 2, 4, 8, 16, and 32 for feature extraction, generating four feature layers: C2, C3, C4, and C5.
[0032] The top-down and horizontal connection parts of the feature pyramid are set as follows: the C5 feature layer first undergoes 1x1 convolution to obtain the M5 feature layer; the M5 feature layer is first upsampled by 2 times, and then the C4 feature layer is horizontally connected by 1*1 convolution to obtain the M4 feature layer; the M4 feature layer is first upsampled by 2 times, and then the C3 feature layer is horizontally connected by 1*1 convolution to obtain the M3 feature layer; the M3 feature layer is first upsampled by 2 times, and then the C2 feature layer is horizontally connected by 1*1 convolution to obtain the M2 feature layer; the M2 feature layer, M3 feature layer, M4 feature layer, and M5 feature layer are respectively subjected to 3*3 convolution to obtain the final P2 feature layer, P3 feature layer, P4 feature layer, and P5 feature layer; among them, in the top-down part, the M2 feature layer, M3 feature layer, M4 feature layer, and M5 feature layer are horizontally connected and then embedded in the dynamic receptive field module.
[0033] The setting of the dynamic receptive field module comprises the following steps:
[0034] The small receptive field branch is set to pass through a 3*3 convolution kernel first, the number of input and output channels remains unchanged, and the size of the feature map is maintained by padding. Then it is processed in two ways: one way is to perform global average pooling to generate a 1*1*c vector, which is used to describe the global features at a deep level and generate activation factors for each dimension. The other way is to perform global maximum pooling to obtain deep spatial information and generate corresponding activation factors. The activation factors of the two ways have the same dimension, which are directly added and then fed back to each channel through the nonlinear activation functions relu and sigmoid.
[0035] y1=relu(bn(c 1 (x)))
[0036] Among them, x is the feature layer after the feature layers of the feature pyramid M are horizontally connected, and c 1represents 3*3 convolution, relu is the linear rectification function, bn is the batch normalization layer, and y1 is the output layer;
[0037] A1=f2(relu(f1(avgpool(y1))))
[0038] B1=f2(relu(f1(maxpool(y1))))
[0039] Among them, f1 and f2 are fully connected layers, avgpool and maxpool represent global average pooling and global maximum pooling respectively, A1 is the one-dimensional vector output by global average pooling, and B1 is the one-dimensional vector output by global maximum pooling;
[0040] The large receptive field branch is set to use the dilated convolution branch to increase the receptive field, and the 3*3 convolution kernel with an expansion coefficient of 2 is used to improve the feature extraction effect. The number of input and output channels remains unchanged, and the size of the feature map is maintained by padding. Then, the activation factor is obtained through global average pooling and global maximum pooling and fed back to each channel. The expression is as follows:
[0041] y2=relu(bn(c 2 (x)))
[0042] Among them, x is the feature layer after the feature layers of the feature pyramid M are horizontally connected, and c 2 represents a 3*3 convolution with a dilation coefficient of 2, relu is a linear rectification function, bn is a batch normalization layer, and y2 is the output layer;
[0043] A2=f2(relu(f1(avgpool(y2))))
[0044] B2=f2(relu(f1(maxpool(y2))))
[0045] Among them, f1 and f2 are fully connected layers, avgpool and maxpool represent global average pooling and global maximum pooling respectively, A2 is the one-dimensional vector output by global average pooling, and B2 is the one-dimensional vector output by global maximum pooling;
[0046] f(x)=σ(A1+B1)*y1+σ(A2+B2)*y2,
[0047] Among them, relu and σ are linear rectification function, linear activation function and sigmoid function respectively, A1, B1, y1 and A2, B2, y2 are two branches activated by different feature layers respectively, and f(x) is the output feature layer of the two branches with large and small receptive fields;
[0048] The dilated convolution is used to control the branches to use different receptive fields. The small receptive field branch and the small receptive field branch are respectively subjected to global average pooling and global maximum pooling to generate corresponding activation factors, thereby fully integrating the context information and the information between each channel, and finally feeding back to the feature extraction network;
[0049] output=(input+2*padding-dilation*(kernel-1)-1) / stride+1,
[0050] Among them, padding is the filling coefficient, dilation is the expansion coefficient, kernel is the convolution kernel size, stride is the step size, and input is the size of the input feature layer.
[0051] Beneficial Effects
[0052] Compared with the prior art, the invention discloses a power transmission line pin defect detection method based on dynamic receptive field. Different receptive fields are adaptively used in the fusion process of different layers of feature pyramid network, and the context information of multiple channels is fully integrated. Receptive fields of different sizes and information within channels are utilized. The network is generated by spatially activating regions, and information acquisition of regions of interest is enhanced, feature extraction of deep convolutional networks is improved, and more information is retained for the final classification and regression of the detector, thereby further improving the accuracy and recognition rate of pin defect detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a method sequence diagram of the present invention;
[0054] Figure 2 A close-up of a pin defect of a power transmission line detected by the method of the present invention;
[0055] Figure 3 This is a distant view of pin defects in a power transmission line detected by the method of the present invention. DETAILED DESCRIPTION
[0056] In order to have a further understanding and recognition of the structural features and the effects achieved by the present invention, a preferred embodiment and accompanying drawings are used for detailed description as follows:
[0057] like Figure 1 As shown, a method for detecting pin defects in a power transmission line based on a dynamic receptive field according to the present invention comprises the following steps:
[0058] The first step is to obtain and preprocess the transmission line pin defect samples: obtain the transmission line pin defect samples, and preprocess the defect samples by data cleaning and data enhancement.
[0059] In practical applications, a large number of original pin defect samples obtained by the State Grid Transmission Panorama Platform can be cleaned first to remove some low-quality defect images. For the images of a small number of pin defect types, data augmentation can be performed by rotation, contrast ratio change, size scaling, etc. When making the training data set, the pin defect category balance is controlled to avoid too many or too few defect types affecting the training effect of the model.
[0060] The second step is to build a transmission line pin defect detection model: a transmission line pin defect detection model consisting of a feature extraction network embedded with a dynamic receptive field module to fuse context and inter-channel information, a region generation network embedded with a spatial activation module, and a defect detection network. The detection model using a dynamic receptive field can adapt to the size changes of transmission line pin defects caused by shooting angles in actual scenarios. Large-sized defects are more advantageous with a large receptive field, while small-sized defects are more advantageous with a small receptive field. The traditional feature extraction network convolution layer only uses a single receptive field, making it difficult to fuse receptive fields of different sizes in the feature extraction network.
[0061] The specific steps are as follows:
[0062] (1) Based on the feature pyramid, a feature extraction network is constructed that combines dynamic receptive fields of different sizes and integrates contextual information.
[0063] A1) Setting of feature pyramid: Setting the feature pyramid includes three parts: bottom-up, top-down and horizontal connection.
[0064] The setting of feature pyramid includes the following steps:
[0065] A11) setting the bottom-up part of the feature pyramid to pass through 5 convolution layers, the 5 convolution layers corresponding to the convolutions with step sizes of 2, 4, 8, 16, and 32 perform feature extraction in sequence, and generate four feature layers: C2 feature layer, C3 feature layer, C4 feature layer, and C5 feature layer;
[0066] A12) The top-down and horizontal connection parts of the feature pyramid are set, and the C5 feature layer is first subjected to 1x1 convolution to obtain the M5 feature layer; the M5 feature layer is first upsampled by 2 times, and then the C4 feature layer is added and horizontally connected by 1*1 convolution to obtain the M4 feature layer; the M4 feature layer is first upsampled by 2 times, and then the C3 feature layer is added and horizontally connected by 1*1 convolution to obtain the M3 feature layer; the M3 feature layer is first upsampled by 2 times, and then the C2 feature layer is added and horizontally connected by 1*1 convolution to obtain the M2 feature layer; the M2 feature layer, the M3 feature layer, the M4 feature layer, and the M5 feature layer are respectively subjected to 3*3 convolution to obtain the final P2 feature layer, the P3 feature layer, the P4 feature layer, and the P5 feature layer; wherein, in the top-down part, the M2 feature layer, the M3 feature layer, the M4 feature layer, and the M5 feature layer are horizontally connected and then embedded in a dynamic receptive field module to further improve the network's extraction of feature information.
[0067] A2) Setting the dynamic receptive field module: The dynamic receptive field module is set to consist of two branches, namely, a large receptive field branch and a small receptive field branch with receptive fields of different sizes and integrating context and channel information. As the network depth increases, the receptive field also increases. However, the size of sample defects in actual scenarios varies greatly, and the receptive field is often too large or too small, which is not conducive to the model learning the characteristic information of pin defects.
[0068] The setting of the dynamic receptive field module comprises the following steps:
[0069] A21) Set the small receptive field branch to first pass through a 3*3 convolution kernel, the number of input and output channels remains unchanged, and the size of the feature map is maintained by padding. Then it is processed in two ways: one way is to perform global average pooling to generate a 1*1*c vector, which is used to describe the global features at a deep level and generate activation factors for each dimension. The other way is to perform global maximum pooling to obtain deep spatial information and generate corresponding activation factors. The activation factors of the two ways have the same dimension, which are directly added and then fed back to each channel through nonlinear activation functions relu and sigmoid;
[0070] y1=relu(bn(c 1 (x)))
[0071] Among them, x is the feature layer after the feature layers of the feature pyramid M are horizontally connected, and c 1 represents 3*3 convolution, relu is the linear rectification function, bn is the batch normalization layer, and y1 is the output layer;
[0072] A1=f2(relu(f1(avgpool(y1))))
[0073] B1=f2(relu(f1(maxpool(y1))))
[0074] Among them, f1 and f2 are fully connected layers, avgpool and maxpool represent global average pooling and global maximum pooling respectively, A1 is the one-dimensional vector output by global average pooling, and B1 is the one-dimensional vector output by global maximum pooling;
[0075] A22) Set the large receptive field branch to use the dilated convolution branch to increase the receptive field, use a 3*3 convolution kernel with an expansion coefficient of 2 to improve the feature extraction effect, keep the number of input and output channels unchanged, and use padding to keep the size of the feature map. Then, the activation factor is obtained through global average pooling and global maximum pooling and fed back to each channel. The expression is as follows:
[0076] y2=relu(bn(c 2 (x)))
[0077] Among them, x is the feature layer after the feature layers of the feature pyramid M are horizontally connected, and c 2 represents a 3*3 convolution with a dilation coefficient of 2, relu is a linear rectification function, bn is a batch normalization layer, and y2 is the output layer;
[0078] A2=f2(relu(f1(avgpool(y2))))
[0079] B2=f2(relu(f1(maxpool(y2))))
[0080] Among them, f1 and f2 are fully connected layers, avgpool and maxpool represent global average pooling and global maximum pooling respectively, A2 is the one-dimensional vector output by global average pooling, and B2 is the one-dimensional vector output by global maximum pooling;
[0081] f(x)=σ(A1+B1)*y1+σ(A2+B2)*y2,
[0082] Among them, relu and σ are linear rectification function, linear activation function and sigmoid function respectively, A1, B1, y1 and A2, B2, y2 are two branches activated by different feature layers respectively, and f(x) is the output feature layer of the two branches with large and small receptive fields;
[0083] A23) The dilated convolution control branch uses different receptive fields. The small receptive field branch and the small receptive field branch are respectively subjected to global average pooling and global maximum pooling to generate corresponding activation factors to fully integrate the context information and the information between each channel, and finally feedback to the feature extraction network;
[0084] output=(input+2*padding-dilation*(kernel-1)-1) / stride+1,
[0085] Among them, padding is the filling coefficient, dilation is the expansion coefficient, kernel is the convolution kernel size, stride is the step size, and input is the size of the input feature layer.
[0086] (2) By activating the spatial information of the region of interest, a spatially activated region generation network is constructed.
[0087] B1) Setting of the region generation network: The region generation network is set to consist of two branches, a 1*1*18 convolution classification branch using softmax loss, and a 1*1*36 convolution regression branch using smooth-L1 loss. The region of interest is generated through the two branches;
[0088] B2) The spatial activation module is set to consist of two branches, which are placed after each region of interest to further activate the spatial information of these regions. Specifically, one branch averages the pixel values of the corresponding points of the feature maps of all channels to obtain a feature map, and the global information of the feature layer is better obtained by averaging. The other branch takes the maximum value of the pixel values of the corresponding points of the feature maps of all channels to obtain another feature map, and the texture information of the feature layer is better obtained by taking the maximum value. Then, the two feature maps are spliced together, and then activated by a 3*3 convolution layer and a sigmoid nonlinear function to generate a heat map and feedback it to the corresponding region of interest;
[0089] f(x)=σ{c[mean(x)+max(x)]}*x,
[0090] Among them, c is a 3*3 convolution, σ is a sigmoid activation function, x is the input region of interest, f(x) is the region of interest after activation, max(x) is the maximum value function, and mean(x) is the average value function.
[0091] (3) Constructing a defect detection network: The defect detection network makes the final prediction of the region of interest after spatial activation, which includes two branches: a classification branch with a 1*1*C convolution using softmax loss and a localization branch with a 1*1*4 convolution using smooth-L1 loss.
[0092] The third step is to train the transmission line pin defect detection model: the preprocessed transmission line pin defect samples are sent to the transmission line pin defect detection model for training. During the training process, as the model network depth gradually increases, more semantic information will be learned, but the feature layer resolution is gradually reduced due to multiple pooling, and part of the pixel information is lost. Therefore, spatially activating these feature layers during the training process can improve the model's learning of defect features, while the common spatial activation method consumes a lot of computing resources.
[0093] The specific steps are as follows:
[0094] (1) The power transmission line pin defect detection model is trained for 24 rounds. The training dataset image H*W is input into the feature extraction network, where H is the image height and W is the image width.
[0095] C1) The training dataset image H*W enters the bottom-up part of the feature extraction network: the input image consists of 5 convolutional layers, each of which consists of 64-channel 1*1 convolution, 64-channel 3*3 convolution and 256-channel 1*1 convolution, generating C2 feature layer, C3 feature layer, C4 feature layer, C5 feature layer respectively;
[0096] C2) Then enter the top-down and lateral connection part of the feature extraction network:
[0097] Specifically, the C5 feature layer first undergoes a 1*1 convolution to obtain the M5 feature layer; the M5 feature layer is first upsampled by 2 times, and then the C4 feature layer is horizontally connected with a 1*1 convolution to obtain the M4 feature layer; the M4 feature layer is first upsampled by 2 times, and then the C3 feature layer is horizontally connected with a 1*1 convolution to obtain the M3 feature layer; the M3 feature layer is first upsampled by 2 times, and then the C2 feature layer is horizontally connected with a 1*1 convolution to obtain the M2 feature layer; the M2 feature layer, M3 feature layer, M4 feature layer, and M5 feature layer are then subjected to 3*3 convolutions respectively to obtain the final P2 feature layer, P3 feature layer, P4 feature layer, and P5 feature layer.
[0098] (2) The feature extraction network obtains output layers P2, P3, P4, and P5, which are input into the region generation network of the embedded spatial activation module to generate regions of interest on the four feature layers P2, P3, P4, and P5, respectively.
[0099] (3) The region of interest obtained by the region generation network embedded in the spatial activation module is cut out from the P2 feature layer, P3 feature layer, P4 feature layer, and P5 feature layer through a pooling layer and input into the defect detection network. The fully connected branch of the defect detection network uses the cross entropy loss, and the bounding box regression branch uses the GIOU loss as the loss function.
[0100] The fourth step is to obtain the image of the pins of the transmission line to be detected: obtain the image of the pins of the transmission line to be detected and perform preprocessing.
[0101] The fifth step is to obtain the transmission line pin defect detection results: the preprocessed transmission line pin image to be detected is input into the trained transmission line pin defect detection model to obtain the transmission line pin defect detection results.
[0102] like Figure 2 and Figure 3 As shown in the figure, the trained model achieved good results in the scenarios of overhead lines in mountainous areas and transmission lines above cities, and the detection recall rate reached 78%.
[0103] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions only describe the principles of the present invention. The present invention may be subject to various changes and improvements without departing from the spirit and scope of the present invention. These changes and improvements fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the attached claims and their equivalents.
Claims
1. A method for detecting pin defects in power transmission lines based on dynamic receptive field. It is characterized in that The following steps are involved: 11) Acquisition and preprocessing of transmission line pin defect samples: Acquire transmission line pin defect samples and perform data cleaning and data enhancement preprocessing on the defect samples; 12) Construction of a transmission line pin defect detection model: A transmission line pin defect detection model is constructed, which consists of a feature extraction network embedded in a dynamic receptive field module to fuse context and inter-channel information, a region generation network embedded in a spatial activation module, and a defect detection network; The construction of the transmission line pin defect detection model includes the following steps: 121) Based on the feature pyramid, a feature extraction network is constructed that combines dynamic receptive fields of different sizes and integrates context information; 1211) Setting of feature pyramid: Setting the feature pyramid includes three parts: bottom-up, top-down and horizontal connection; 1212) Setting a dynamic receptive field module: Setting the dynamic receptive field module to consist of two branches, which are a large receptive field branch and a small receptive field branch with receptive fields of different sizes and integrating context and channel information; 122) By activating the spatial information of the region of interest, a spatially activated region generation network is constructed; 1221) Setting of region generation network: The region generation network is set to consist of two branches, a 1*1*18 convolution classification branch using softmax loss, and a 1*1*36 convolution regression branch using smooth-L1 loss. The region of interest is generated through the two branches; 1222) The spatial activation module is set to consist of two branches, which are placed after each region of interest to further activate the spatial information of these regions. Specifically, one branch averages the pixel values of the corresponding points of the feature maps of all channels to obtain a feature map, and the global information of the feature layer is better obtained by averaging. The other branch takes the maximum value of the pixel values of the corresponding points of the feature maps of all channels to obtain another feature map, and the texture information of the feature layer is better obtained by taking the maximum value. Then, the two feature maps are spliced together, and then activated by a 3*3 convolution layer and a sigmoid nonlinear function to generate a heat map to feedback to the corresponding region of interest; f(x)=σ{c[mean(x)+max(x)]}*x, Among them, c is 3*3 convolution, σ is sigmoid activation function, x is the input region of interest, f(x) is the region of interest after activation, max(x) is the maximum value function, and mean(x) is the average value function; 123) Construct defect detection network: The defect detection network makes the final prediction of the region of interest after spatial activation, including two branches, a 1*1*C convolution classification branch using softmax loss, and a 1*1*4 convolution positioning branch using smooth-L1 loss; 13) Training of the transmission line pin defect detection model: sending the preprocessed transmission line pin defect samples to the transmission line pin defect detection model for training; 14) Obtaining an image of pins on a power transmission line to be detected: obtaining an image of pins on a power transmission line to be detected and performing preprocessing; 15) Obtaining the transmission line pin defect detection results: The preprocessed transmission line pin image to be detected is input into the trained transmission line pin defect detection model to obtain the transmission line pin defect detection results.
2. According to claim 1, a method for detecting pin defects in power transmission lines based on dynamic receptive fields, It is characterized in that The training of the transmission line pin defect detection model includes the following steps: 21) Set the power transmission line pin defect detection model to be trained for 24 rounds, and input the training data set image H*W into the feature extraction network, where H is the image height and W is the image width; 211) The training dataset image H*W enters the bottom-up part of the feature extraction network: the input image consists of 5 convolutional layers, each of which consists of 64-channel 1*1 convolution, 64-channel 3*3 convolution and 256-channel 1*1 convolution, generating four feature layers: C2 feature layer, C3 feature layer, C4 feature layer and C5 feature layer respectively; 212) Then enter the top-down and lateral connection parts of the feature extraction network: Specifically, the C5 feature layer is first subjected to 1*1 convolution to obtain the M5 feature layer; M5 is first upsampled by 2 times, and then the C4 feature layer is connected horizontally by 1*1 convolution to obtain the M4 feature layer; the M4 feature layer is first upsampled by 2 times, and then the C3 feature layer is connected horizontally by 1*1 convolution to obtain the M3 feature layer; the M3 feature layer is first upsampled by 2 times, and then the C2 feature layer is connected horizontally by 1*1 convolution to obtain the M2 feature layer; the M2 feature layer, M3 feature layer, M4 feature layer, and M5 feature layer are then subjected to 3*3 convolution respectively to obtain the final P2 feature layer, P3 feature layer, P4 feature layer, and P5 feature layer; 22) The P2 feature layer, P3 feature layer, P4 feature layer, and P5 feature layer output by the feature extraction network are input into the region generation network of the embedded spatial activation module to generate regions of interest on the four feature layers of P2 feature layer, P3 feature layer, P4 feature layer, and P5 feature layer respectively; 23) The region of interest obtained by the region generation network embedded in the spatial activation module is cut out from the P2 feature layer, P3 feature layer, P4 feature layer, and P5 feature layer through a pooling layer and input into the defect detection network. The fully connected branch of the defect detection network uses the cross entropy loss, and the bounding box regression branch uses the GIOU loss as the loss function.
3. According to claim 1, a method for detecting pin defects in power transmission lines based on dynamic receptive fields, It is characterized in that The setting of the feature pyramid includes the following steps: 31) Set the bottom-up part of the feature pyramid to pass through 5 convolution layers, and the 5 convolution layers correspond to the convolutions with step sizes of 2, 4, 8, 16, and 32 for feature extraction, generating four feature layers: C2 feature layer, C3 feature layer, C4 feature layer, and C5 feature layer; 32) Set the top-down and horizontal connection parts of the feature pyramid, that is, the C5 feature layer first undergoes 1x1 convolution to obtain the M5 feature layer; the M5 feature layer is first upsampled by 2 times, and then the C4 feature layer is added and horizontally connected by 1*1 convolution to obtain the M4 feature layer; the M4 feature layer is first upsampled by 2 times, and then the C3 feature layer is added and horizontally connected by 1*1 convolution to obtain the M3 feature layer; the M3 feature layer is first upsampled by 2 times, and then the C2 feature layer is added and horizontally connected by 1*1 convolution to obtain the M2 feature layer; the M2 feature layer, the M3 feature layer, the M4 feature layer, and the M5 feature layer are respectively subjected to 3*3 convolution to obtain the final P2 feature layer, P3 feature layer, P4 feature layer, and P5 feature layer; wherein, in the top-down part, the M2 feature layer, the M3 feature layer, the M4 feature layer, and the M5 feature layer are horizontally connected and then embedded in the dynamic receptive field module.
4. According to claim 1, a method for detecting pin defects in power transmission lines based on dynamic receptive fields, It is characterized in that The setting of the dynamic receptive field module comprises the following steps: 41) Set the small receptive field branch to first pass through a 3*3 convolution kernel, the number of input and output channels remains unchanged, and the size of the feature map is maintained by padding. Then it is processed in two ways: one way is to perform global average pooling to generate a 1*1*c vector, which is used to describe the global features at a deep level and generate activation factors for each dimension. The other way is to perform global maximum pooling to obtain deep spatial information and generate corresponding activation factors. The activation factors of the two ways have the same dimension, which are directly added and then fed back to each channel through nonlinear activation functions relu and sigmoid; y1=relu(bn(c 1 (x))) Among them, x is the feature layer after the feature layers of the feature pyramid M are horizontally connected, and c 1 represents 3*3 convolution, relu is the linear rectification function, bn is the batch normalization layer, and y1 is the output layer; A1=f2(relu(f1(avgpool(y1)))) B1=f2(relu(f1(maxpool(y1)))) Among them, f1 and f2 are fully connected layers, avgpool and maxpool represent global average pooling and global maximum pooling respectively, A1 is the one-dimensional vector output by global average pooling, and B1 is the one-dimensional vector output by global maximum pooling; 42) Set the large receptive field branch to use the dilated convolution branch to increase the receptive field, use a 3*3 convolution kernel with an expansion coefficient of 2 to improve the feature extraction effect, keep the number of input and output channels unchanged, and use padding to keep the size of the feature map. Then, the activation factor is obtained through global average pooling and global maximum pooling and fed back to each channel. The expression is as follows: y2=relu(bn(c 2 (x))) Among them, x is the feature layer after the feature layers of the feature pyramid M are horizontally connected, and c 2 represents a 3*3 convolution with a dilation coefficient of 2, relu is a linear rectification function, bn is a batch normalization layer, and y2 is the output layer; A2=f2(relu(f1(avgpool(y2)))) B2=f2(relu(f1(maxpool(y2)))) Among them, f1 and f2 are fully connected layers, avgpool and maxpool represent global average pooling and global maximum pooling respectively, A2 is the one-dimensional vector output by global average pooling, and B2 is the one-dimensional vector output by global maximum pooling; f(x)=σ(A1+B1)*y1+σ(A2+B2)*y2, Among them, relu and σ are linear rectification function, linear activation function and sigmoid function respectively, A1, B1, y1 and A2, B2, y2 are two branches activated by different feature layers respectively, and f(x) is the output feature layer of the two branches with large and small receptive fields; 43) The dilated convolution control branch uses different receptive fields. The small receptive field branch and the small receptive field branch are respectively subjected to global average pooling and global maximum pooling to generate corresponding activation factors to fully integrate the context information and the information between each channel, and finally feedback to the feature extraction network; output=(input+2*padding-dilation*(kernel-1)-1) / stride+1, Among them, padding is the filling coefficient, dilation is the expansion coefficient, kernel is the convolution kernel size, stride is the step size, and input is the size of the input feature layer.
Citation Information
Patent Citations
Tunnel surface defect detection method based on feature pyramid network
CN111524117A
Deformable convolution-based power transmission tower pin defect detection method
CN114155246A