Defect Location System and Method for Transmission Line Insulator String Based on Cascade Detection Strategy
By introducing a cascade detection strategy and an insulator string candidate area positioning network in the SOLOv2 network, the problem of low detection accuracy and speed of insulator string defects under complex environments and limited computing resources is solved, and efficient and accurate detection results are achieved.
Patent Information
- Application Number
- CN202510147175.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-11
AI Technical Summary
In the case of complex environments, long-tail data and computing resources, the SOLOv2 network has poor detection accuracy and low speed in insulator string defect detection.
A transmission line insulator string defect positioning system based on cascade detection strategy is adopted, combined with the insulator string candidate area positioning network and the improved SOLOv2 network, efficient identification and defect detection are achieved through a two-stage detection strategy.
It significantly improves the accuracy and speed of insulator string defect detection, especially in complex backgrounds, and meets the efficiency requirements of real-time monitoring of transmission lines.
Smart Images

Figure CN119600031B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, relates to equipment monitoring in power engineering, and particularly relates to a transmission line insulator string defect location system and method based on a cascade detection strategy. Background Art
[0002] With the rapid development of the automation and intelligence of the power system, the intelligent detection of power lines has become a research hotspot in the field of power inspection. Traditional manual inspection and detection methods based on image processing, although they can provide certain support in some scenarios, have problems such as low detection efficiency, poor accuracy, and being greatly affected by environmental conditions. Especially in complex environments, traditional methods often have difficulty dealing with complex power equipment and diverse environmental changes.
[0003] In this context, computer vision and deep learning technologies have been widely applied to the intelligent inspection of power lines, especially object detection and instance segmentation technologies. The SOLOv2 network, as an advanced instance segmentation model, performs well in many visual detection tasks. Through a unique structure design and an efficient learning mechanism, the SOLOv2 network can effectively detect and segment targets in relatively complex scenarios and has strong performance.
[0004] However, although the SOLOv2 network has achieved good results in many object detection tasks, it still faces some challenges. First, when dealing with dense targets or small objects, the SOLOv2 network may be affected by blurred boundaries, resulting in a decrease in segmentation accuracy. Second, the processing ability of the SOLOv2 network for long-tail distribution data still needs to be optimized. Especially when some categories in the dataset are relatively scarce, the model may show biases or omissions. In addition, the computational complexity of the SOLOv2 network is relatively high, and when deployed on an embedded platform with limited resources, it may face greater computational and storage pressures, which limits its application in real-time monitoring.
[0005] Therefore, although the SOLOv2 network has strong advantages in many object detection tasks, it still has certain limitations in dealing with complex environments, long-tail data, and limited computational resources. To address these problems, combining the advantages of the traditional SOLOv2 network to improve the robustness, efficiency, and accuracy of the model has become an important research direction. Summary of the Invention
[0006] The purpose of the present invention is to aim at the deficiencies in the above-mentioned prior art and provide a transmission line insulator string defect location system and method based on a cascade detection strategy to achieve efficient identification of insulator strings on transmission lines and perform defect detection, thereby solving the technical problems of poor accuracy and low speed in insulator string defect detection in the prior art.
[0007] To achieve the above object, the present invention adopts the following technical solutions to implement.
[0008] The transmission line insulator string defect location system based on a cascaded detection strategy provided by the present invention includes:
[0009] An insulator string candidate region location network for identifying and annotating a bounding box in an image to be processed to locate the insulator string candidate region;
[0010] An instance segmentation network for defect identification and instance segmentation of the insulator string candidate region; the instance segmentation network uses an improved SOLOv2 network, including a backbone sub-network, a neck sub-network, and a prediction sub-network; the backbone sub-network is used to extract multi-scale multi-layer features from the input image features; the neck sub-network is used to fuse multi-scale multi-layer features; the neck sub-network includes a first fusion unit and a second fusion unit; the first fusion unit, in the direction of increasing scale, performs upsampling and addition processing on the output features of the backbone sub-network to obtain the multi-layer fusion output of the first fusion unit; the second fusion unit, in the direction of decreasing scale, performs aggregation and addition processing based on a bidirectional attention mechanism on the fusion outputs of each layer of the first fusion unit to obtain the multi-layer fusion output of the second fusion unit; the prediction sub-network is used to identify the defects of the insulator string and perform instance segmentation based on the multi-layer fusion output of the second fusion unit.
[0011] In an implementable manner, the insulator string candidate region location network serves as a cascaded window before the improved SOLOv2 network to first identify the insulator string from the image to be processed. In the present invention, the insulator string candidate region location network uses the YOLOv9 network.
[0012] In an implementable manner, the backbone sub-network of the instance segmentation network uses an improved ResNet network; the improved ResNet network includes a plurality of residual modules arranged in sequence; the residual modules have the same structure, including a plurality of convolutional modules and a deformable convolutional layer arranged in sequence, and the input features input to the residual module are sequentially passed through the plurality of convolutional modules and the deformable convolutional layer and then concatenated with the input features by channel to obtain the output features of the residual module.
[0013] In one implementable manner, the backbone sub-network further includes a CQSFM module disposed before a plurality of residual modules. The CQSFM module includes two-dimensional convolutional layers of four different scales, a max pooling layer, and a splicing unit. First, the candidate regions of the insulator strings in the input image are convolved by the two-dimensional convolutional layers of four different scales, and then passed through the max pooling layer to reduce the size to half of the original. Then, the four convolutional outputs are spliced in the clockwise direction to form a new candidate region of the insulator string, which is used as the input of the residual module. By using convolutional kernels of different sizes to extract features, the problem of the insulator string being too large or too small caused by the change of the UAV shooting position can be effectively addressed.
[0014] In one implementable manner, in the neck sub-network of the instance segmentation network, the second fusion unit includes a plurality of BiELAN modules. Specifically, the number of BiELAN modules is 1 less than the number of layers of the fusion output features of the first fusion unit. The BiELAN module is used to aggregate the input features based on the bidirectional attention mechanism. The specific operation is as follows: the input feature image is divided into several blocks at a set step size, and three different blocks are randomly selected to construct a key matrix, a query matrix, and a value matrix respectively. First, the key matrix and the query matrix are multiplied and then processed by the Softmax function, and the calculation result is multiplied by the value matrix, and the obtained result is aggregated by the ELAN module; repeating this step for the input feature image to obtain the corresponding aggregation result, that is, the output features of the BiELAN module.
[0015] In one implementable manner, the prediction sub-network includes a classification branch, a mask kernel branch, and a mask feature branch. The classification branch is used to identify the category of the insulator string (such as normal or defective) based on the fusion result of the neck sub-network. The mask kernel branch is used to generate the dynamic convolutional kernel weights for segmenting the insulator string defects based on the fusion result of the neck sub-network. The mask feature branch is used to generate the instance segmentation image of the insulator string defects based on the fusion result of the neck sub-network. The mask feature branch performs dynamic convolution with the outputs of the classification branch and the mask kernel branch to obtain the final instance segmentation result of the insulator string defect region. The classification branch, the mask kernel branch, and the mask feature branch all adopt the conventional structures disclosed in the traditional SOLOv2.
[0016] The present invention also provides a method for locating insulator string defects on transmission lines based on a cascaded detection strategy, including the following steps:
[0017] S1. Construct a dataset for instance segmentation of insulator strings on transmission lines;
[0018] S2. Use the dataset for instance segmentation of insulator strings on transmission lines to train the insulator string defect location system for transmission lines;
[0019] S3. Input the collected transmission line images containing insulator strings into the trained defect localization system for insulator strings of transmission lines to obtain the instance segmentation results of insulator string defects.
[0020] The above step S1 includes the following sub-steps:
[0021] S11. Collect a number of real image samples containing insulator strings and perform instance segmentation annotation on the insulator strings.
[0022] S12. Simulate different environmental conditions through virtual scene building software to perform simulation processing on the real image samples containing insulator strings to obtain a number of virtual image samples containing insulator strings.
[0023] S13. The virtual image samples and real image samples constitute the instance segmentation dataset for insulator strings of transmission lines.
[0024] In the above step S11, use a drone equipped with a high-resolution image acquisition device to conduct aerial photography of the transmission line regularly to obtain multiple images containing insulator strings on the transmission line and its installation accessories. The flexibility and efficiency of the drone ensure the comprehensiveness and real-time nature of data collection, thus providing sufficient image data for subsequent defect detection.
[0025] Perform instance segmentation annotation on the collected images containing insulator strings to obtain an accurate real dataset for instance segmentation of insulator strings and their defects. And through manual or semi-automatic annotation tools, accurately label the boundaries of each insulator string and its defects to ensure the high quality of the dataset and provide a reliable basis for model training.
[0026] In the above step S12, use virtual scene building software (such as Unreal Tournament 4 (UE4)) to simulate and build the overall structure of the transmission line to generate a large number of high-quality virtual datasets. Through simulation technology, the states of insulator strings under different environmental conditions can be simulated, such as different weather, lighting, angles, etc.
[0027] It is also possible to perform diversified enhancement processing on the virtual datasets, including operations such as rotation, scaling, color jitter, etc., to expand the real dataset and improve the adaptability of the model to diversified environmental conditions.
[0028] The above step S2 includes the following sub-steps:
[0029] S21. Divide the instance segmentation dataset for insulator strings of transmission lines into a training set and a test set.
[0030] S22. Use the training set to train the defect localization system for insulator strings of transmission lines; the above step S22 includes the following sub-steps:
[0031] S221. Input the samples in the training set into the transmission line insulator string defect localization system. First, locate the candidate regions of the insulator string through the insulator string candidate region localization network and label them. Then, input the image with the labeled candidate regions of the insulator string into the instance segmentation network to obtain the insulator string defect instance segmentation image.
[0032] S222. Calculate the loss value based on the predicted insulator string defect instance segmentation image and the labeled instance segmentation image of each sample.
[0033] S223. Optimize the parameters of the transmission line insulator string defect localization system according to the loss value.
[0034] Repeat the above steps S221 - S223 until the transmission line insulator string defect localization system converges.
[0035] S23. Use the test set to test and verify the trained transmission line insulator string defect localization system.
[0036] The above step S3 includes the following sub - steps:
[0037] S31. Input the collected transmission line image containing the insulator string into the insulator string candidate region localization network to locate the candidate regions of the insulator string and label them.
[0038] S32. Input the image with the located candidate regions of the insulator string into the instance segmentation network to obtain the insulator string defect instance segmentation image.
[0039] The transmission line insulator string defect localization system and method based on the cascade detection strategy provided by the present invention have the following beneficial effects:
[0040] (1) The present invention first locates the insulator string through the insulator string candidate region localization network, and then performs instance segmentation on the insulator string defects through the instance segmentation network, forming a two - stage cascade detection strategy, which can more accurately locate the insulator string defects.
[0041] (2) The present invention sets the insulator string candidate region localization network as a cascade window in front of the improved SOLOv2 network, fully combining the capabilities of fast object detection and the improved SOLOv2 network in high - precision instance segmentation. This combination not only significantly improves the detection accuracy of insulator string defects, especially performs well in complex backgrounds and small targets, but also greatly improves the overall detection speed, meeting the high - efficiency requirements of real - time monitoring of transmission lines. Moreover, by using the bounding box information of the cascade window as the input of the improved SOLOv2 network, the system can accurately define the specific region of instance segmentation, avoiding redundant calculations in the whole - image range, significantly reducing the waste of computing resources, and improving resource utilization.
[0042] (3) In the backbone subnet of the improved SOLOv2 network of the present invention, a BiELAN module (Bidirectional Enhanced Local Attention Network) is introduced. Through a double-layer routing attention mechanism and an adaptive query mechanism, more efficient multi-scale feature fusion is achieved, enhancing the recognition ability for insulator strings of different sizes. In addition, deformable convolution (DCN) is integrated into the residual structure of the improved SOLOv2 network, enabling the model to dynamically adjust the sampling positions of convolutional kernels, flexibly adapting to various geometric transformations and deformations of insulator strings in images, and further improving the segmentation quality and accuracy.
[0043] (4) In the backbone subnet of the improved SOLOv2 network of the present invention, a clockwise four-scale fusion module is introduced. By extracting features using convolutional kernels of different sizes, the problem of insulator strings being too large or too small caused by changes in the UAV shooting position can be effectively addressed.
[0044] (5) The present invention uses virtual scene construction software to generate a large amount of high-quality virtual data, and combines diversified data augmentation techniques (such as rotation, scaling, color jitter, etc.) to effectively expand the real dataset. This method not only solves the problems of insufficient training data and lack of diversity, improving the adaptability of the model under different environmental conditions, but also enhances the robustness and reliability of the system in the face of various complex environments and different types of insulator string defects through multiple technical improvements (such as BiELAN, multi-step convolution, deformable convolution, etc.), ensuring stability and efficiency in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 Schematic diagram of the structure of a transmission line insulator string defect location system based on a cascade detection strategy provided in Embodiment 1 of the present invention;
[0046] Figure 2 Schematic diagram of the principle of the CQSFM module;
[0047] Figure 3 Schematic diagram of the structure of the residual module;
[0048] Figure 4 Schematic diagram of the principle of deformable convolution;
[0049] Figure 5 Schematic diagram of the principle of the BiELAN module;
[0050] Figure 6 Schematic diagram of the structure of the ELAN module;
[0051] Figure 7 Schematic diagram of the flow of a method for locating transmission line insulator string defects based on a cascade detection strategy provided in Embodiment 2 of the present invention;
[0052] Figure 8 This is an example of the image of the transmission line insulator string collected in Embodiment 2 of the present invention;
[0053] Figure 9 is Figure 8 the output image after the image of the transmission line insulator string in Detailed implementation manners
[0054] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other implementation manners obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope protected by the present invention.
[0055] Embodiment 1
[0056] This embodiment provides a transmission line insulator string defect localization system based on a cascade detection strategy, which includes an insulator string candidate region localization network and an instance segmentation network. The insulator string candidate region localization network is used to identify the insulator string in the image to be processed and mark the bounding box to locate the insulator string candidate region. The instance segmentation network is used to perform defect identification and instance segmentation on the insulator string candidate region.
[0057] The insulator string candidate region localization network, as a cascaded window division before the improved SOLOv2 network, first calibrates the insulator string from the image to be processed. In this embodiment, the YOLOv9 network is used for the insulator string candidate region localization network. For the network structure, refer to "YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information" published by Chien-Yao Wang et al. (Chien-Yao Wang, I-Hau Yeh, and Hong-Yuan Mark Liao, "YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information", arXiv:2402.13616, February 29, 2024). The YOLOv9 network is used to perform object detection on the input image, which can quickly locate the candidate regions of the insulator string and obtain the bounding box coordinates of each candidate region. With its efficient detection ability, the YOLOv9 network can provide accurate bounding box information while maintaining high speed. The bounding box coordinates output by the YOLOv9 network are used as the input of the improved SOLOv2 network to accurately define the region for instance segmentation of the improved SOLOv2 network; moreover, the insulator string to be subjected to instance segmentation can be determined according to the area of the insulator string, so as to further improve the instance segmentation accuracy of the insulator string defects. In this way, the improved SOLOv2 network does not need to perform segmentation within the entire image range, thereby improving the segmentation efficiency and accuracy and reducing the waste of computing resources.
[0058] Therefore, by first obtaining the insulator string through cascaded window division and combining the speed advantage of the YOLOv9 network; then introducing the high-precision instance segmentation ability of the improved SOLOv2 network and adopting a two-stage cascaded detection strategy, the ability to capture small targets and detailed features is significantly improved. This strategy not only speeds up the detection speed but also improves the detection accuracy, especially showing excellent performance in the detection of insulator string defects in complex backgrounds.
[0059] In this embodiment, the improved SOLOv2 network is used for the instance segmentation network, which includes a backbone sub-network, a neck sub-network, and a prediction sub-network.
[0060] (1) Backbone sub-network
[0061] The backbone sub-network is used to extract multi-layer features of different scales from the input image features.
[0062] Such as Figure 1As shown, in this embodiment, the backbone sub-network uses an improved ResNet50 network. The improved ResNet50 network includes a CQSFM module (Clockwise Quad-Scale Fusion Module) and four residual modules (Residual Module 1 - Residual Module 4) arranged in sequence.
[0063] As Figure 2 shown, the CQSFM module includes two-dimensional convolutional layers of four different scales, a max pooling layer, and a concatenation unit. First, the input image is convolved using two-dimensional convolutional layers of four different scales (defined as Conv1, Conv2, Conv3, and Conv4 in this embodiment, with kernel sizes of 1×1, 3×3, 5×5, and 7×7 respectively), and then passed through a 2×2 max pooling layer (MaxPooling) to reduce the size to half of the original (both the length and width are reduced to half of the original); then, the four convolutional outputs are concatenated in a clockwise direction to form a new image, which is used as the input to the input layer. By using convolutional kernels of different sizes to extract features, the problem of the insulator string being too large or too small due to changes in the UAV shooting position can be effectively addressed.
[0064] The residual modules have the same structure. As Figure 3 shown, it includes a 1×1 convolutional module, a 3×3 convolutional module, a 1×1 convolutional module, and a deformable convolutional layer arranged in sequence. Each convolutional module includes a convolutional layer, a batch normalization layer, and a ReLU activation function. The input features input to the residual module pass through each convolutional module and the deformable convolutional layer in sequence and are concatenated with the input features by channel to obtain the output features of the residual module. Through multi-step convolutional operations, the spatial variation information of the input feature map is obtained. Moreover, multi-step convolution can capture richer spatial context information (i.e., spatial variations of different scales and directions), generate offsets for adjusting the sampling positions of the deformable convolution, and provide a more accurate basis for adjusting the sampling positions for the subsequent deformable convolution.
[0065] The deformable convolutional layer (Deformable Convolution, DCN) is an innovative convolutional neural network structure that allows the convolutional kernel to dynamically adjust its sampling position according to the content in the input feature map during image processing. DCN controls the sampling points of the convolutional kernel by introducing offsets, enabling the model to more flexibly adapt to various shape and size changes. The principle of DCN is as Figure 4As shown, when performing deformable convolution operations, the original convolution kernel first determines an initial sampling grid on the input feature map; then, adjusts the position of this grid according to the learned offsets to obtain the corresponding offset region. In this way, even in the face of complex deformations such as rotation, scaling, or distortion, the convolution kernel can effectively capture the key features of the target object. This mechanism not only improves the quality and accuracy of instance segmentation but also enhances the model's adaptability to target objects with different shapes and angles. In this embodiment, deformable convolution is integrated into the residual module of SOLOv2. The feature map generated in the previous step is used as the input. The offsets are determined according to the extracted spatial variation information, and the sampling positions of the convolution kernel on the feature map are controlled using the offsets, thereby realizing that the deformable convolution dynamically adjusts the sampling positions of the convolution kernel. Due to the existence of the offsets, the output feature map can better reflect the true shape and boundaries of the target objects in the input image, which is crucial for improving the performance of semantic segmentation and instance segmentation tasks; especially when dealing with insulator strings with different shapes and angles, it can maintain a high level of segmentation quality and accuracy; and improve the model's adaptability to complex deformations.
[0066] In this embodiment, the output features C2 - C5 of Residual Module 1 - Residual Module 4 are used as the output of the backbone sub-network.
[0067] (2) Neck Sub-network
[0068] The neck sub-network is used to fuse multi-level features of different scales. The neck sub-network includes a first fusion unit and a second fusion unit.
[0069] The first fusion unit, in the direction of increasing scale, performs upsampling and addition operations on the output features of the backbone sub-network to obtain the multi-level fusion output of the first fusion unit. The first fusion unit includes three upsampling modules, namely Upsampling Module 1, Upsampling Module 2, and Upsampling Module 3. The first fusion unit takes the output features C2 - C5 of the backbone sub-network as the input, and obtains the fused output features M2 - M5 of the first fusion unit through each sampling module and horizontal connection. Specifically, the feature C5 is processed by a 1×1 convolution (Conv1×1) to obtain the feature M5. The result of the feature M5 processed by Upsampling Module 1 is added to the result of the feature C4 processed by a 1×1 convolution (here it refers to element-wise addition) to obtain the feature M4. The result of the feature M4 processed by Upsampling Module 2 is added to the result of the feature C3 processed by a 1×1 convolution to obtain the feature M3. The result of the feature M3 processed by Upsampling Module 3 is added to the result of the feature C2 processed by a 1×1 convolution to obtain the feature M2.
[0070] The second fusion unit processes the fusion outputs of each layer of the first fusion unit in the direction of decreasing scale through aggregation and addition based on a bidirectional attention mechanism to obtain the multi-layer fusion output of the second fusion unit. The second fusion unit includes three BiELAN modules, namely BiELAN module 1, BiELAN module 2, and BiELAN module 3. The second fusion unit takes the output features M2 - M5 of the first fusion unit as input, and through each BiELAN module and horizontal connection, obtains the fusion output features P2 - P5 of the second fusion unit. Specifically, the feature M2 is processed by a 3×3 convolution (Conv3×3) to obtain the feature P2, the result of processing the feature P2 by BiELAN module 1 is added to the result of processing the feature M3 by a 3×3 convolution to obtain the feature P3, the result of processing the feature P3 by BiELAN module 2 is added to the result of processing the feature M4 by a 3×3 convolution to obtain the feature P4, and the result of processing the feature P4 by BiELAN module 3 is added to the result of processing the feature M5 by a 3×3 convolution to obtain the feature P5.
[0071] In this embodiment, the BiELAN (Bidirectional Level Efficient Layer Aggregation Networks) module is introduced into the improved SOLOv2 network to enhance the fusion effect of multi-scale features. The BiELAN module aggregates the input features based on a bidirectional attention mechanism. Through bidirectional information flow and the attention mechanism, it can more effectively integrate feature information of different scales and improve the model's detection ability for multi-scale targets.
[0072] As Figure 5 shown, the specific operation of the BiELAN module is as follows: According to a set stride (for example s =2), the input feature image is divided into several blocks, and three different blocks are randomly selected to construct the key matrix K , the query matrix Q , and the value matrix V , that is, the image blocks with a stride of 1 are stacked; for the key matrix K and the value matrix V , a coefficient k is also given (in this embodiment, k =2, that is, the image blocks are stacked twice), so the dimensions of the key matrix K and the value matrix V are , the dimension of the query matrix is , W , H represent the width and height of the input feature image, C represents the number of channels. First, use the key matrix K and the query matrix QMultiply and then process through the Softmax function (i.e., mm&Softmax in the figure, where mm represents matrix multiplication), and then multiply the calculation result with the value matrix V Multiply, and the obtained result is aggregated by the ELAN module to obtain the aggregated result of the three image patches. Repeating this step for all image patches of the input feature image gives the aggregated result of the input feature image (without repeated selection of each image patch), which is also the output feature of the BiELAN module.
[0073] The above ELAN (Efficient Layer Aggregation Networks) module structure is as Figure 6 shown. It includes two branches. The first branch includes a 1×1 convolution module, and the second branch includes a 1×1 convolution module and four 3×3 convolution modules arranged in sequence. The input features of the ELAN module are processed through the two branches respectively. The output feature of the 1×1 convolution module in the second branch and the output feature of the second 3×3 convolution module are residually connected to the output of the fourth 3×3 convolution module, and then added to the output feature obtained from the first branch and then processed through a 1×1 convolution module to obtain the output result of the ELAN module.
[0074] The BiELAN module adopts a two-layer routing attention mechanism to filter out the most irrelevant key-value pairs from the insulator string features extracted at the coarse-grained level. This mechanism highlights key information and suppresses irrelevant information by dynamically adjusting the attention weights, thereby improving the effectiveness and robustness of feature representation. And through the adaptive query mechanism, content awareness in the sparse mode is achieved. This mechanism dynamically adjusts the query strategy according to the content of the input features, enabling the model to capture key features more flexibly and improving the intelligent level of feature fusion.
[0075] Moreover, when applying token-to-token attention, the BiELAN module used in this embodiment avoids relying on computationally intensive sparse matrix algorithms that merge memory operations, and instead uses a hardware-friendly dense matrix algorithm to collect key-value tokens. This optimization significantly reduces the computational complexity, improves the execution efficiency of the attention mechanism, and enables the model to have better real-time performance while maintaining high performance.
[0076] (3) Prediction sub-network
[0077] The prediction sub-network is used to identify and perform instance segmentation on the defects of the insulator string based on the multi-layer fusion output of the second fusion unit.
[0078] The prediction sub-network includes a classification branch, a mask kernel branch, and a mask feature branch; the classification branch is used to identify the insulator string category (such as normal or defective) based on the fusion result of the neck sub-network; the mask kernel branch is used to generate dynamic convolution kernel weights for segmenting insulator string defects based on the fusion result of the neck sub-network; the mask feature branch is used to generate an instance segmentation image of the insulator string defect based on the fusion result of the neck sub-network; the output of the mask feature branch and the classification branch and the mask kernel branch are dynamically convolved to obtain the final instance segmentation result of the insulator string defect area. The classification branch, the mask kernel branch, and the mask feature branch all adopt the conventional structures disclosed in the traditional SOLOv2. For details, please refer to "SOLOv2: Dynamic and Fast Instance Segmentation" published in the 34th Conference on Neural Information Processing Systems, by Xinlong Wang, Rufeng Zhang, Tao Kong, Lei Li, and Chunhua Shen (Xinlong Wang, Rufeng Zhang, Tao Kong, Lei Li, Chunhua Shen, "SOLOv2:Dynamic and Fast Instance Segmentation", 34th Conference on Neural InformationProcessing Systems (NeurIPS 2020), Vancouver, Canada).
[0079] Embodiment 2
[0080] This embodiment provides a method for locating insulator string defects on a transmission line based on a cascaded detection strategy, as Figure 7 shown, including the following steps:
[0081] S1. Construct an instance segmentation dataset for the insulator string on the transmission line.
[0082] This step is to construct an instance segmentation dataset for the insulator string on the transmission line for training the insulator string defect location system on the transmission line based on the cascaded detection strategy provided in Embodiment 1.
[0083] This step includes the following sub-steps:
[0084] S11. Collect a number of real image samples containing insulator strings and perform instance segmentation annotation on the insulator strings.
[0085] In this step, a DJI Matrice 300 RTK drone equipped with a DJI Zenmuse P1 high-resolution camera is used for image collection, as Figure 8As shown in the figure. The drone conducts regular aerial photography of the target transmission line once a week, and the flight altitude is set at 120 meters to ensure that the acquired images have sufficient resolution and coverage. Each aerial photography generates images that cover the insulator strings on the transmission line and its installation fittings, ensuring the comprehensiveness and real-time nature of the data. During the image acquisition process, GPS marking and the drone flight control system are used to record the geographical location information of each image, facilitating subsequent data processing and model training. In this embodiment, the number of image samples containing insulator strings collected is 1000.
[0086] The collected insulator string images are labeled for instance segmentation using the LabelMe annotation tool. The insulator strings in each image are boundary-labeled using an accurate polygon tool to generate high-quality instance segmentation masks. During the annotation process, ensure that the boundaries of each insulator string are clear and accurate, avoiding overlap and omission. After the annotation is completed, all the annotated data undergoes quality inspection to ensure the high accuracy and consistency of the dataset, providing a reliable real data basis for subsequent model training.
[0087] S12. Simulate different environmental conditions through virtual scene construction software, and perform simulation processing on the real image samples containing insulator strings to obtain several virtual image samples containing insulator strings.
[0088] In this embodiment, the Unreal Engine 4 (UE4) is used to build a virtual transmission line environment based on the three-dimensional model of the real transmission line. In the virtual environment, images of insulator strings under various conditions such as different weather (e.g., sunny, rainy, foggy), lighting (day-night changes), or / and angles (different shooting angles) are generated. Based on the 1000 real images obtained previously, the number of virtual images generated through the Unreal Engine 4 reaches 5000.
[0089] This step also performs diverse enhancement processing through the scripts built into the Unreal Engine 4, including operations such as random rotation (±30 degrees), scaling (0.8 times to 1.2 times), color jitter (changes in brightness, contrast, saturation), and adding noise, to perform data enhancement processing on the generated virtual images.
[0090] S13. The virtual image samples and the real image samples constitute the instance segmentation dataset of the transmission line insulator strings.
[0091] The enhanced virtual image samples are combined with the real image samples to form the instance segmentation dataset of the transmission line insulator strings, further expanding the dataset scale and improving the generalization ability and environmental adaptability of the model.
[0092] In this embodiment, the real image samples and the virtual image samples are combined to form an instance segmentation dataset of the transmission line insulator strings with a total of 6000 images.
[0093] S2. Use the instance segmentation dataset of transmission line insulator strings to train the transmission line insulator string defect localization system.
[0094] This step S2 includes the following sub-steps:
[0095] S21. Divide the instance segmentation dataset of transmission line insulator strings into a training set and a test set.
[0096] In this step, the instance segmentation dataset of transmission line insulator strings is randomly divided into a training set (4800 images) and a test set (1200 images) according to a ratio of 8:2. The training set contains 4800 images and their corresponding instance segmentation masks, and the test set contains 1200 images and their corresponding instance segmentation masks. This division method ensures that the model can fully learn different types and diverse data features during the training process and can effectively evaluate the performance of the model during the test phase.
[0097] S22. Use the training set to train the transmission line insulator string defect localization system. This step includes the following sub-steps:
[0098] S221. Input the samples in the training set into the transmission line insulator string defect localization system. First, locate and label the insulator string candidate regions through the insulator string candidate region localization network; then input the image with the labeled insulator string candidate regions into the instance segmentation network to obtain the insulator string defect instance segmentation image.
[0099] In this embodiment, the YOLOv9 network pre-trained using a public dataset is used as the insulator string candidate region localization network. Perform object detection on each image in the training set and the test set. The YOLOv9 network is configured with an input resolution of 1024×1024 pixels and achieves efficient object detection through GPU acceleration (NVIDIA RTX 3090). YOLOv9 outputs the bounding box coordinates ( x , y , width, height) of each insulator string and its confidence score. The detection results are stored in JSON format for subsequent processing.
[0100] Use the bounding box coordinates detected by the YOLOv9 network as the input region of the instance segmentation network (improved SOLOv2 network). The improved SOLOv2 network only performs instance segmentation operations within these candidate regions, avoiding redundant calculations across the entire image. Specifically, use a Python script to convert the bounding box coordinates output by the YOLOv9 network into the input format required by SOLOv2 and pass it to the improved SOLOv2 network for precise segmentation to obtain the insulator string defect instance segmentation image.
[0101] S222. Calculate the loss value based on the instance segmentation images of the insulator string defects predicted from each sample and the labeled instance segmentation images.
[0102] In this step, for the instance segmentation network part, the overall loss function L ins mainly includes:
[0103] Classification branch loss L cate : Determine the category to which the current instance (insulator string / defect) belongs (such as "defect" or "normal");
[0104] Mask branch loss L mask : Generate a binary segmentation mask for the insulator string (or defect area);
[0105] Therefore, this part can usually be expressed as:
[0106] (1);
[0107] Wherein, , respectively represent weighting factors.
[0108] (2);
[0109] Wherein, p t represents p when it is a positive sample (i.e., the network predicts the category correctly), and p 1 - p when it is a negative sample (i.e., the network predicts the category incorrectly), γ represents the category probability obtained by the network; α t is called the focusing factor, which can suppress easy-to-separate samples and enhance the gradient weight of difficult samples;
[0110] (3);
[0111] Wherein, represents the mask branch loss calculated based on X and Y ; X represents the image mask matrix generated by the network, x i ∈[0,1] represents the mask pixel value generated by the network; Y represents the true image mask matrix,y i ∈[0,1] represents the true mask pixel value.
[0112] Based on the insulator string defect instance segmentation images predicted for each sample and the annotated instance segmentation images, the corresponding loss values are calculated through the above overall loss function.
[0113] S223. Optimize the parameters of the transmission line insulator string defect localization system according to the loss value.
[0114] Based on the loss value calculated in step S222, the parameters of the improved SOLOv2 network are optimized through the gradient descent optimization algorithm Adam.
[0115] Repeat the above steps S221 - S223 until the transmission line insulator string defect localization system converges. When the loss value is less than the set threshold, it can be considered that the improved SOLOv2 network converges, and the trained transmission line insulator string defect localization system is obtained.
[0116] S23. Use the test set to test and verify the trained transmission line insulator string defect localization system.
[0117] In this step, the trained transmission line insulator string defect localization system is tested using the test set, and the test effect is evaluated. For example, the system is evaluated through indicators such as accuracy, intersection over union (IoU), and false positive rate (FPR). If the evaluation effect does not meet the set requirements, the network parameters of the system need to be adjusted, and the system is retrained using the training set until the set requirements are met.
[0118] S3. Input the collected transmission line images containing insulator strings into the trained transmission line insulator string defect localization system to obtain the insulator string defect instance segmentation result.
[0119] This step S3 includes the following sub - steps:
[0120] S31. Input the collected transmission line images containing insulator strings into the insulator string candidate region localization network to locate and annotate the insulator string candidate regions.
[0121] S32. Input the image with the located insulator string candidate regions into the instance segmentation network to obtain the insulator string defect instance segmentation image.
[0122] Through the above steps S31 and S32, insulator strings and their defect images with different colors highlighted are generated, such as Figure 9As shown in the figure; the identified insulator string is marked by a red frame in the figure; and the normal area and the defective area of the specified insulator string in the figure are effectively instance-segmented and distinguished by different colors.
[0123] In summary, the present invention adopts a two-stage cascaded detection strategy. First, the insulator string candidate area positioning network quickly locates the insulator string candidate area, and then the instance segmentation network performs high-precision instance segmentation within the limited area. This strategy not only speeds up the overall detection speed but also improves the ability to capture small targets and detailed features, especially performing well in the detection of insulator string defects in complex backgrounds. The test results show that this strategy increases the detection speed by 30% and improves the detection accuracy by 15% at the same time.
[0124] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A transmission line insulator string defect location system based on cascade detection strategy, characterized in that: include: An insulator string candidate region positioning network is used to identify the insulator strings in the processed image and mark the bounding box to locate the insulator string candidate region; An instance segmentation network is used to perform defect recognition and instance segmentation on candidate areas of insulator strings; the instance segmentation network uses an improved SOLOv2 network, including a trunk subnetwork, a neck subnetwork, and a prediction subnetwork; The backbone sub-network is used to extract multi-layer features of different scales from input image features; The backbone subnetwork of the instance segmentation network uses an improved ResNet network; the improved ResNet network includes a plurality of residual modules arranged in sequence; the residual modules have the same structure, including a plurality of convolution modules and a deformable convolution layer arranged in sequence, and the input features of the input residual modules are sequentially spliced with the input features according to channels after passing through a plurality of convolution modules and a variable convolution layer to obtain the output features of the residual modules; the backbone subnetwork also includes a CQSFM module arranged in front of the plurality of residual modules, and the CQSFM module includes four two-dimensional convolution layers of different scales, a maximum pooling layer and a splicing unit; first, the candidate region of the insulator string in the input image is convolved with four two-dimensional convolution layers of different scales, and then the maximum pooling layer is passed to reduce the size to half of the original size, and the length and width are both reduced to half of the original size; then, the four convolution outputs are spliced in a clockwise direction to form a new candidate region of the insulator string as the input of the residual module; The neck subnetwork is used to fuse multi-layer features of different scales; the neck subnetwork includes a first fusion unit and a second fusion unit; the first fusion unit performs upsampling and addition processing on each output feature of the trunk subnetwork in the direction of increasing scale to obtain the multi-layer fusion output of the first fusion unit; the second fusion unit performs aggregation and addition processing on each layer of fusion output of the first fusion unit in the direction of reducing scale based on a bidirectional attention mechanism to obtain the multi-layer fusion output of the second fusion unit; the second fusion unit includes a plurality of BiELAN modules, and the BiELAN module is used to aggregate the input features based on the bidirectional attention mechanism, and the specific operation is as follows: according to the set step size, the input feature image is divided into several blocks, three different blocks are randomly selected to respectively construct a key matrix, a query matrix and a value matrix, firstly, the key matrix and the query matrix are multiplied and then processed by the Softmax function, and the calculated result is multiplied by the value matrix, and the result is aggregated by the ELAN module; repeating this step for the input feature image obtains the corresponding aggregation result, that is, the output feature of the BiELAN module; The prediction subnetwork is used to identify defects of the insulator string and perform instance segmentation based on the multi-layer fusion output of the second fusion unit.
2. The transmission line insulator string defect location system based on cascade detection strategy according to claim 1 is characterized in that: The insulator string candidate area positioning network uses the YOLOv9 network.
3. The transmission line insulator string defect location system based on cascade detection strategy according to claim 1 is characterized in that: The specific operation of the BiELAN module is as follows: according to the set step size, the input feature image is divided into several blocks, and three different blocks are randomly selected to construct the key matrix K , query matrix Q Sum Matrix V ; For the bond matrix K Sum Matrix V , and also given the coefficient k , so the key matrix K Sum Matrix V The dimension is , the dimension of the query matrix is , W , H represents the width and height of the input feature image, C Indicates the number of channels; first use the key matrix K and the query matrix Q Multiply and process by Softmax function, and then add the result to the value matrix V The results are multiplied and aggregated by the ELAN module to obtain the aggregation results of the three image blocks; this step is repeated for all image blocks of the input feature image to obtain the aggregation result of the input feature image.
4. The transmission line insulator string defect location system based on cascade detection strategy according to claim 1 is characterized in that: The prediction subnetwork includes a classification branch, a mask core branch and a mask feature branch; The classification branch is used to identify the insulator string category based on the fusion result of the neck sub-network; The mask kernel branch is used to generate dynamic convolution kernel weights for segmenting insulator string defects according to the fusion result of the neck sub-network; The mask feature branch is used to generate an insulator string instance segmentation image based on the fusion result of the neck sub-network; the mask feature branch is dynamically convolved with the output of the classification branch and the mask kernel branch to obtain the final insulator string defect area instance segmentation result.
5. A method for locating defects in insulator strings of transmission lines based on a cascade detection strategy, characterized in that: The following steps are involved: S1. Construct a data set for instance segmentation of insulator strings for transmission lines; S2. Training the transmission line insulator string defect location system according to any one of claims 1 to 4 using a transmission line insulator string instance segmentation dataset; S3. Input the collected transmission line image containing the insulator string into the trained transmission line insulator string defect location system to obtain the insulator string defect instance segmentation result.
6. The method for locating defects in insulator strings of power transmission lines based on a cascade detection strategy according to claim 5, characterized in that: The step S1 comprises the following sub-steps: S11, collecting a number of real image samples containing insulator strings, and performing instance segmentation and annotation on the insulator strings; S12, simulating different environmental conditions through virtual scene building software, performing simulation processing on real image samples containing insulator strings, and obtaining a number of virtual image samples containing insulator strings; S13, virtual image samples and real image samples constitute the transmission line insulator string instance segmentation dataset.
7. The method for locating defects in insulator strings of power transmission lines based on a cascade detection strategy according to claim 5, characterized in that: The step S2 comprises the following sub-steps: S21, dividing the transmission line insulator string instance segmentation dataset into a training set and a test set; S22, using the training set to train the transmission line insulator string defect location system; the step S22 includes the following sub-steps: S221, inputting samples in the training set into the transmission line insulator string defect location system, firstly locating and marking the insulator string candidate area through the insulator string candidate area location network; then inputting the image of the marked insulator string candidate area into the instance segmentation network to obtain the insulator string defect instance segmentation image; S222, calculating the loss value based on the instance segmentation image of the insulator string defect predicted by each sample and the labeled instance segmentation image; S223, optimizing the parameters of the transmission line insulator string defect location system according to the loss value; Repeat the above steps S221-S223 until the transmission line insulator string defect location system converges; S23. Use the test set to test and verify the trained transmission line insulator string defect location system.
8. The method for locating defects in insulator strings of power transmission lines based on a cascade detection strategy according to claim 5, characterized in that: The step S3 comprises the following sub-steps: S31, inputting the collected transmission line image containing the insulator string into the insulator string candidate area positioning network, locating the insulator string candidate area and marking it; S32, inputting the image with the candidate area of the insulator string located into the instance segmentation network to obtain the instance segmentation image of the insulator string defect.
Citation Information
Patent Citations
Insulator fault identification method and system based on target detection and instance segmentation
CN115294473A
Image direction prediction method based on multi-scale fusion and attention mechanism
CN115761258A
Glass insulator lightning stroke discharge defect identification method, device and equipment
CN116338392A
Solar cell defect detection method based on attention convolutional neural network
CN116523820A
Improved SOLO-based unstructured road scene instance segmentation method and system
CN116543358A