A real-time inspection method for damaged insulators of power transmission lines
By constructing a high-resolution insulator damage detection model, real-time inspection is carried out using KASEN-FreSAN self-attention bilateral context feature network and YOLO Head for real-time inspection, the problems of inspection lag and identification detection accuracy in the existing technology are solved, and more accurate and robust target recognition is achieved.
Patent Information
- Application Number
- CN202510174332.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The existing drone inspection technology has problems with patrol lag and identification detection accuracy when identifying damaged insulators on transmission lines.
A real-time inspection method is adopted to build a high-resolution insulator damage detection model, including a backbone feature extraction network, a KASEN-FreSAN self-attention bilateral context feature network and a enhanced feature extraction network. By obtaining transmission line images in real time, extracting keyframes, performing feature extraction and enhancement, and finally using YOLO Head for accurate damage detection.
It improves the accuracy and robustness of target recognition, enhances feature extraction capabilities, and improves the detection performance of the model in complex scenarios. It is suitable for real-time intelligent patrol tasks of drones with damaged insulators in high-resolution transmission line.
Smart Images

Figure CN119649310B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned aerial vehicle inspection technology, and in particular to a real-time inspection method for damaged insulators of power transmission lines. Background Art
[0002] Insulators are an important component of power transmission lines. They are responsible for supporting the conductors and providing insulation to ensure the safe transmission of electric energy. The performance of insulators is directly related to the safety and reliability of power transmission lines. However, due to long-term exposure to harsh natural environments, insulators may suffer damage, such as broken sheds, missing spring pins, and even breakage. If these problems are not detected and repaired in time, they may cause line failures or power loss, affecting power supply. The working mode of existing power inspection drones is mainly based on image recognition algorithms, supplemented by manual inspections. Although the recognition algorithm greatly reduces the manual burden, the inspection lag problem and the recognition detection accuracy problem are still two major problems in the industry. Summary of the invention
[0003] In view of the above-mentioned technical problems to be solved, the present invention provides a real-time inspection method for damaged insulators of transmission lines, which can enhance the accuracy and robustness of target recognition.
[0004] In order to solve the above technical problems, the technical solution proposed by the present invention is:
[0005] A real-time inspection method for damaged insulators of a power transmission line comprises the following steps:
[0006] Step S1, constructing a high-resolution insulator damage detection model, which includes: a backbone feature extraction network, a KASEN-FreSAN self-attention bilateral context feature network, and an enhanced feature extraction network;
[0007] Step S2, acquiring the image of the transmission line to be operated in real time, extracting key frames from the real-time image, and inputting the key frames into the damage detection model;
[0008] Step S3, the backbone feature extraction network in the damage detection model performs feature extraction with downsampling operation on the key frame of the transmission line image to be operated to obtain global feature information;
[0009] Step S4, feature enhancement based on KASEN-FreSAN self-attention bilateral context feature network;
[0010] Step S5, strengthening the feature extraction network to further extract and fuse multi-scale features of feature information;
[0011] Step S6, YOLO Head takes the three sets of fused features as input and generates the final accurate damage detection output;
[0012] Step S7, performing structural reparameterization processing on the convolutional layer in the backbone feature extraction network, simplifying the complex multi-branch structure in the training phase into a single efficient structure in the inference phase.
[0013] As a further improvement of the above technical solution:
[0014] Preferably, in step S1, the backbone feature extraction network is composed of multiple layers of convolutional neural layers, and multiple groups of effective feature layers are obtained through multiple groups of multi-branch stacking-transition modules. The KASEN-FreSAN self-attention bilateral context feature network is located after the backbone feature extraction network; the KASEN-FreSAN self-attention bilateral context feature network includes a KASEN attention module, a FreSAN attention module and a BFM module; the enhanced feature extraction network has three feature modeling channel branches and a feature aggregation layer.
[0015] Preferably, the transition module is provided with a left branch and a right branch, the left branch is a maximum pooling with a stride of 2x2 and a 1x1 convolution, the right branch is a 1x1 convolution and a convolution with a convolution kernel size of 3x3 and a stride of 2x2, and the results of the two branches are stacked again at output.
[0016] Preferably, the multi-branch stacking module contains two branches, the first branch is a convolutional normalization activation function to change the number of channels; the second branch first passes through a 1*1 convolution module to change the number of channels, and then passes through four 3*3 convolution modules to perform feature extraction to form four feature layers; after stacking, the four feature layers are again subjected to a convolutional normalization activation function for feature integration.
[0017] Preferably, in step S3, an image of size H*W*C is input, and the image is subjected to a convolutional layer for preliminary feature extraction to obtain a preliminary feature map of half the initial size, and C is further enriched by the initial input in the channel dimension; then a set of convolutional normalization activation functions and multi-branch stacking modules are passed through to reduce the size of the feature map and increase the number of channels; three sets of multi-branch stacking-transition modules are passed through to obtain three sets of effective feature layers, one set of which is used as the input of the KASEN-FreSAN self-attention bilateral context feature network, and the other two sets of effective feature layers are used as the input of the enhanced feature extraction network.
[0018] Preferably, the step S4 specifically includes the following contents:
[0019] S4-1, in the left branch, first input the feature map with the dimension of H*W*C, use KAN convolution to perform preliminary feature extraction, obtain the new feature map Y, apply the global average pooling operation to the feature map Y after KAN convolution; compress the spatial dimension H*W of Y to obtain the feature map Z; learn through the FC fully connected layer operation to obtain the channel attention weight vector A; multiply the channel attention weight vector A and the feature map Y after KAN convolution channel by channel;
[0020] S4-2, in the right branch, use the two-dimensional discrete Fourier transform to obtain the gradient vector; select the gain coefficient to obtain the enhanced frequency domain gradient, use the multiplication update method to update the frequency domain image to obtain the enhanced frequency domain image, and use the inverse Fourier transform to convert it back to the spatial domain; then construct the spatial attention mask and multiply it with the original input image to obtain the final output result;
[0021] S4-3, in the BFM module, the feature maps of the left branch and the right branch are mixed to generate a mixed feature map to complete the fusion of the bilateral network.
[0022] Preferably, in step S5, the three groups of effective feature layers are fused; in the enhanced feature extraction network, the effective feature layers that have been obtained are used to continue to extract features; at the same time, using the Panet structure, the features are upsampled to achieve feature fusion, and the features are downsampled again to achieve feature fusion.
[0023] The real-time inspection method for damaged insulators of power transmission lines provided by the present invention has the following advantages compared with the prior art:
[0024] The real-time inspection method for damaged insulators of power transmission lines of the present invention adopts the KASEN-FreSAN self-attention bilateral context feature network. In actual shooting, it is found through image data statistics that there are problems of various shooting angles in the spatial scale. Insulators with high aspect ratios have equal probability in each angle direction when shooting, there is strong correlation in space, and the insulator materials are different. Materials such as porcelain and glass have large channel differences under illumination changes, colors and different shooting conditions, showing certain differences, which increases the difficulty of detection. The present invention introduces an attention mechanism based on the KASEN-FreSAN bilateral, uses the FreSAN attention mechanism to capture the spatial global information of the insulator, uses the KASEN attention mechanism to improve the detection ability of the model to cope with different shooting conditions, improves the target detection performance of the YOLOv7 model in drone images of complex scenes, and the ability to focus on key feature areas and spectral information, thereby enhancing the accuracy and robustness of target recognition. The present invention has the advantages of efficient feature extraction, multi-scale feature modeling, fast and accurate defect detection and high reasoning efficiency, and is suitable for the real-time intelligent inspection task of damaged drones of high-resolution power transmission line insulators. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a framework diagram of the damage detection model of the present invention.
[0026] Figure 2 It is a framework diagram of the transition module of the present invention.
[0027] Figure 3 It is a framework diagram of the multi-branch stacking module of the present invention.
[0028] Figure 4 is a schematic diagram of the BFM module of the present invention. DETAILED DESCRIPTION
[0029] The specific embodiments of the present invention are described in detail below. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0030] like Figure 1 As shown, the real-time inspection method for damaged insulators of a power transmission line of the present invention comprises the following steps:
[0031] Step S1, construct a high-resolution insulator damage detection model, which includes: a reparameterizable backbone feature extraction network, a weighted KASEN-FreSAN self-attention bilateral context feature network, and an enhanced feature extraction network.
[0032] The backbone feature extraction network is composed of multiple layers of convolutional neural layers. Multiple groups of effective feature layers are obtained through multiple groups of multi-branch stacking-transition modules. The KASEN-FreSAN self-attention bilateral context feature network is located after the backbone feature extraction network.
[0033] In this embodiment, Figure 2 As shown in the figure, three groups of multi-branch stacking-transition modules are used for illustration. The transition module performs downsampling. The two transition modules use a convolution with a convolution kernel size of 3x3 and a stride of 2x2 or a maximum pooling with a stride of 2x2. The two transition modules are combined. The transition module has two branches. The left branch is a maximum pooling with a stride of 2x2 and a 1x1 convolution. The right branch is a 1x1 convolution and a convolution with a convolution kernel size of 3x3 and a stride of 2x2. The results of the two branches are stacked again at the output. Figure 3 As shown in the figure, the multi-branch stacking module contains two branches. The first branch is a convolutional normalization activation function to change the number of channels. The second branch first passes through a 1*1 convolutional module to change the number of channels, and then passes through four 3*3 convolutional modules for feature extraction to form four feature layers. After the four feature layers are stacked, a convolutional normalization activation function is performed again for feature integration.
[0034] The weighted-mergeable KASEN-FreSAN self-attention bilateral contextual feature network consists of three parts: the KASEN attention module, the FreSAN attention module and the BFM module.
[0035] The enhanced feature extraction network has three feature modeling channel branches and a feature aggregation layer.
[0036] Step S2, acquiring the image of the transmission line to be operated in real time, extracting key frames from the real-time image, and inputting it into the constructed damage detection model.
[0037] Generally, transmission line images are acquired through drones.
[0038] Set two keyframes for input and The height and width are equal and are set to H and W respectively. The number of channels C is 3 and they store RGB information respectively. and The difference is greater than the preset range value, Updated to And input the damage detection model.
[0039] The drone patrol process is a video image. Compared with the timed extraction of frames and feeding them into the model detection, the key frame extraction can reduce unnecessary computing costs and avoid repeated identification of insulator defects in the same geographical location. If the previous frame is different from the next frame, it can be determined that the video image has reached a new insulator segment.
[0040] Step S3, the backbone feature extraction network in the damage detection model performs feature extraction with downsampling operation on the key frames of the transmission line image to be operated, so as to gradually obtain high-level, abstract and global feature information.
[0041] In this embodiment, for an image with an input size of H*W*C, the image will go through four convolutional layers: Extract preliminary features and obtain a preliminary feature map of size 160*160*128 The channel dimension of its features C will be further enriched from the initial input of 3 to 128. Then three groups of effective feature layers are obtained through three groups of multi-branch stacking-transition modules, one of which is used as the input of the KASEN-FreSAN self-attention bilateral context feature network, and the other two groups are used as the input of the enhanced feature extraction network.
[0042] Step S4, feature enhancement is performed based on the KASEN-FreSAN self-attention bilateral context feature network. The KASEN-FreSAN self-attention bilateral context feature network includes the KASEN attention module, the FreSAN attention module and the BFM module, which captures and distinguishes the effective features of different insulators, enhances the feature extraction capability and improves the recognition efficiency of the model, so as to improve the overall recognition accuracy of the model.
[0043] KASEN attention module: It uses a learnable nonlinear activation function and a channel attention weight vector to dynamically adjust the output value according to different input data, thereby assigning different weights to different features, suppressing channel features that are not important to the task, reducing the interference of noise and redundant information, and helping to improve classification accuracy.
[0044] FreSAN attention module: Introduces frequency domain enhancement technology and utilizes the high-frequency and low-frequency information of insulator images with high aspect ratios, making the model focus on richer information and improving the accuracy of damaged insulator identification.
[0045] BFM module: Fusion of enhanced features from KASEN attention module and FreSAN attention module.
[0046] Specifically include the following:
[0047] S4-1, the left branch is the KASEN attention module, which performs attention enhancement in the channel space dimension.
[0048] S4-1-1, in the left branch, the feature map with the dimension of H*W*C is first input, and the KAN convolution in the KASEN attention module is used to perform preliminary feature extraction to obtain a new feature map Y with the dimension of H*W*C, where each element , , each element is obtained by applying a learnable nonlinear activation function to the element at the corresponding position in the input feature map and then adding them together.
[0049] S4-1-2, apply global average pooling operation to the feature map Y after KAN convolution.
[0050] Compress the spatial dimension H*W of Y to 1*1, and get a feature map Z with a dimension of 1*1*C. Each channel of this feature map Z It represents the average value of spatial information of the corresponding channel in the feature map Y. The feature map Z is operated and learned through the FC (Fully Connected Layer) to obtain the weight vector A of the channel attention, with a dimension of 1*1*C.
[0051] S4-1-3, multiply the channel attention weight vector A by the feature map Y after KAN convolution channel by channel, so that the KASEN attention module of the left branch can adjust the feature map according to the learned channel importance, enhance the features of important channels, and suppress the features of unimportant channels.
[0052] S4-2, the right branch is the FreSAN attention module, which performs spatial dimension attention enhancement in the frequency domain.
[0053] S4-2-1, use the two-dimensional discrete Fourier transform (2D - DFT) in the right branch to obtain the frequency domain image F(u, v), where u, v are the frequency domain coordinates, and calculate the amplitude spectrum Partial derivatives in the u and v directions give the gradient vector :
[0054] .
[0055] S4-2-2, select a suitable gain coefficient K to obtain the enhanced frequency domain gradient G(u, v) = K , the frequency domain image is updated by multiplication to obtain the enhanced frequency domain image , using the inverse Fourier transform Converting back to the spatial domain yields .by Building a spatial attention mask based on Multiply it with the original input image to get the final output result.
[0056] The appropriate gain coefficient K is obtained by observing the enhanced image and evaluating its visual effect, so as to make the edge of the image clear while not over-enhancing the noise.
[0057] S4-3, the BFM (Bilateral Fusion-Mamba) module consists of eight FSSM modules and is designed as a symmetrical structure. The feature maps of the left branch and the right branch are mixed to generate a mixed feature map to complete the fusion of the bilateral network.
[0058] like Figure 4 As shown, the FSSM module performs the following operations:
[0059] enter: , ,in represents the product of the height and width of the feature map, Indicates the number of channels.
[0060] Output: , the output is a Input feature maps of the same dimension.
[0061] (1) Parameter initialization:
[0062] Get C sets of structured N-by-N matrices from parameters.
[0063] (2) Linear transformation:
[0064] For input Apply a linear transformation ;
[0065] :For input Apply another linear transformation .
[0066] (3) Calculation :
[0067] ,in is a bias vector with a size of .
[0068] (4) Calculation and :
[0069] :calculate ,in, represents tensor product;
[0070] :calculate .
[0071] (5) Feature fusion:
[0072] :Use SSM (State Space Model) to Processing is performed, wherein SSM is implemented by selective scanning.
[0073] (6) Return output:
[0074] Returns the processed feature map .
[0075] The feature maps output by the FreSAN attention module and the KASEN attention module are the input of the BFM module. The two feature maps are recorded as and A method similar to the four-way Mamba module is used to generate two sets of feature maps, as follows:
[0076] ;
[0077] ;
[0078] in, The feature map representing the output of the FreSAN attention module is taken as input, , express After two independent parallel convolutional layers The feature map obtained later is The feature map representing the output of the KASEN attention module is taken as input, , express After two independent parallel convolutional layers The feature map obtained later.
[0079] Next, and The resulting one-dimensional sequence is flattened in four directions (horizontally, vertically, diagonally, and anti-diagonally) and then forwarded to the FSSM module for information integration:
[0080] ;
[0081] in, and They refer to Figure 4 The FSSM modules of the left and right branches of the BFM module, Indicates the direction of flattening. When i=1, it indicates the horizontal direction; when i=2, it indicates the vertical direction; when i=3, it indicates the diagonal direction; when i=4, it indicates the anti-diagonal direction. express A one-dimensional tensor in the i direction, express A one-dimensional tensor in the i direction, represents the flattening in the i direction, Represents the output of the left branch FSSM module in direction i, Represents the output of the right branch FSSM module in direction i, represents the information integration of the left branch FSSM module in the i direction, Represents the information integration of the right branch FSSM module in the i direction.
[0082] After that, the two sets of outputs are processed separately to generate two new feature maps, denoted as and , these graphs are finally combined to form :
[0083] ;
[0084] ;
[0085] ;
[0086] ;
[0087] in, , ,and Indicates generation , and The 1×1 convolutional layer, Indicates restoring the one-dimensional tensor in the i direction to the original feature map shape. express The result of the branch, express The result of the branch.
[0088] The present invention adopts a complete two-branch feature map for multi-scale fusion, the high-level space adopts the channel dimension, and the low-level space adopts the spatial dimension. At the same time, the present invention fuses information more widely and does not accumulate attention in the channel dimension.
[0089] Step S5, strengthening the feature extraction network to further extract multi-scale features from the feature information.
[0090] The three groups of effective feature layers are fused to combine feature information of different scales. In the enhanced feature extraction network, the effective feature layers that have been obtained are used to continue extracting features; at the same time, the Panet (path aggregation network) structure is used to upsample the medium-scale effective feature maps and the feature maps that have passed through the bilateral context feature network to achieve feature fusion; the features after the upsampled features are fused are downsampled again to achieve feature fusion.
[0091] like Figure 1 As shown in , the multi-level scale features correspond to the effective feature layers in the three modules, namely, the multi-branch stacking module (20, 20, 1024), the multi-branch stacking module (40, 40, 1024), and the multi-branch stacking module (80, 80, 512), i.e., the feature layers at smaller, medium, and larger scales.
[0092] In step S6, YOLO Head is used to take the three sets of fused features as input to generate the final accurate damage detection output.
[0093] YOLO Head is the output layer of the YOLOv7 algorithm responsible for the target detection task. It uses the features extracted by the previous network layers to make predictions.
[0094] Step S7, performing structural reparameterization processing on the convolutional layer in the backbone feature extraction network, simplifying the complex multi-branch structure in the training phase into a single efficient structure in the reasoning phase, so as to improve the reasoning efficiency.
[0095] Step S8, use torch.onnx.export function (wrapper function of python learning library) to export the model to onnx format, and transplant and deploy it to yulong810A chip platform. ONNX is a standardized neural network model format, which facilitates model interoperability between different platforms and hardware, and is suitable for real-time reasoning on embedded devices and edge computing devices.
[0096] The present invention provides a real-time inspection method for damaged insulators of power transmission lines, provides a drone inspection method for damaged insulators of power grids, and proposes a real-time high-precision detection method based on YOLOv7-mamba for the existing inspection lag problem and recognition and detection accuracy problem. It includes a space-channel self-attention bilateral context feature network, which establishes a correlation model between image data of the same regional space and channel, indicating the similarity between image data and insulator defect data. Since multi-angle data modeling is involved, more distinguishable and more robust features can be obtained. In addition, considering the real-time requirements, the time complexity of the method is higher. The original complexity based on the Transformer architecture (a deep learning model architecture) is O(n²), and the recognition time is significantly increased when processing high-resolution images. The present method adopts a new SSM sequence model with linear complexity and achieves 5 times the inference throughput.
[0097] The above implementation cases are only preferred embodiments of the present invention and are not intended to limit the present invention in any form. Although the present invention has been disclosed as above with preferred embodiments, they are not intended to limit the present invention. Therefore, any simple modification, equivalent changes and modifications made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention.
Claims
1. A real-time inspection method for damaged insulators of power transmission lines, characterized in that: The following steps are involved: Step S1, constructing a high-resolution insulator damage detection model, which includes: a backbone feature extraction network, a bilateral context feature network based on KASEN-FreSAN self-attention, and an enhanced feature extraction network; Step S2, acquiring the image of the transmission line to be operated in real time, extracting key frames from the real-time image, and inputting the key frames into the damage detection model; Step S3, the backbone feature extraction network in the damage detection model performs feature extraction with downsampling operation on the key frame of the transmission line image to be operated to obtain global feature information; Step S4, feature enhancement based on KASEN-FreSAN self-attention bilateral context feature network; Step S5, strengthening the feature extraction network to further extract and fuse multi-scale features of feature information; Step S6, YOLO Head takes the three sets of fused features as input and generates the final accurate damage detection output; Step S7, performing structural reparameterization processing on the convolutional layer in the backbone feature extraction network, simplifying the complex multi-branch structure in the training phase into a single efficient structure in the inference phase; In step S1, the backbone feature extraction network is composed of multiple layers of convolutional neural layers, and multiple groups of effective feature layers are obtained through multiple groups of multi-branch stacking-transition modules. The KASEN-FreSAN self-attention bilateral context feature network is located after the backbone feature extraction network; the KASEN-FreSAN self-attention bilateral context feature network includes a KASEN attention module, a FreSAN attention module and a BFM module; The enhanced feature extraction network has three feature modeling channel branches and a feature aggregation layer.
2. The real-time inspection method for damaged insulators of a power transmission line according to claim 1 is characterized in that: The transition module is provided with a left branch and a right branch. The left branch is a maximum pooling with a stride of 2x2 and a 1x1 convolution. The right branch is a 1x1 convolution and a convolution with a convolution kernel size of 3x3 and a stride of 2x2. The results of the two branches are stacked again at the output.
3. The real-time inspection method for damaged insulators of a power transmission line according to claim 2 is characterized in that: The multi-branch stacking module contains two branches. The first branch is a convolution normalization activation function to change the number of channels. The second branch first passes through a 1*1 convolution module to change the number of channels, and then passes through four 3*3 convolution modules to perform feature extraction to form four feature layers. After the four feature layers are stacked, a convolutional normalization activation function is performed again for feature integration.
4. The real-time inspection method for damaged insulators of a power transmission line according to claim 3 is characterized in that: In the step S3, an image of size H*W*C is input, and the image is subjected to a convolutional layer for preliminary feature extraction to obtain a preliminary feature map of size 160*160*128, and C is further enriched by the initial input in the channel dimension; and then three groups of effective feature layers are obtained through three groups of multi-branch stacking-transition modules, one group of effective feature layers is used as the input of the KASEN-FreSAN self-attention bilateral context feature network, and the other two groups of effective feature layers are used as the input of the enhanced feature extraction network.
5. The real-time inspection method for damaged insulators of a power transmission line according to claim 4 is characterized in that: The step S4 specifically includes the following contents: S4-1, in the left branch, first input the feature map with the dimension of H*W*C, use KAN convolution to perform preliminary feature extraction, obtain the new feature map Y, apply the global average pooling operation to the feature map Y after KAN convolution; compress the spatial dimension H*W of Y to obtain the feature map Z; learn through the FC fully connected layer operation to obtain the channel attention weight vector A; multiply the channel attention weight vector A and the feature map Y after KAN convolution channel by channel; S4-2, in the right branch, a two-dimensional discrete Fourier transform is used to obtain a gradient vector; a gain coefficient is selected to obtain an enhanced frequency domain gradient, the frequency domain image is updated by a multiplication update method to obtain an enhanced frequency domain image, and an inverse Fourier transform is used to convert it back to the spatial domain; Reconstruct the spatial attention mask and multiply it with the original input image to get the final output result; S4-3, in the BFM module, the feature maps of the left branch and the right branch are mixed to generate a mixed feature map to complete the fusion of the bilateral network.
6. The real-time inspection method for damaged insulators of a power transmission line according to claim 4 is characterized in that: In step S5, the three groups of effective feature layers are fused; in the enhanced feature extraction network, the obtained effective feature layers are used to continue extracting features; At the same time, the Panet structure is used to upsample the features to achieve feature fusion, and the features are downsampled again to achieve feature fusion.
Citation Information
Patent Citations
Insulator defect target detection method based on improved YOLOv4
CN116385335A
Image super-resolution reconstruction method based on lightweight double-branch network
CN116664398A