Power transmission line insulator defect detection method and device based on improved Yolo11n
By improving the feature extraction and loss function of the Yolo11n model, the problem of low recognition accuracy of micro cracks and scintillation defects in the detection of insulator defects in transmission line is solved, and efficient and accurate insulator defect detection is achieved.
Patent Information
- Application Number
- CN202510521630.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art in the detection of insulator defects in transmission line, especially the identification accuracy of micro cracks and scintillation defects is not high, and the computing resources are large, making it difficult to deploy on drones or patrol robots.
By improving the Yolo11n model, the P1, P3 and P5 layers of the backbone feature extraction network are replaced as DSC_ADown module, the ContextGuide FPN module and Dy_SPD module are introduced, and the iEMA attention mechanism and Wasserstein-IoU loss function are combined to enhance the feature perception and detection accuracy of small targets and reduce the calculation amount.
It significantly improves the detection accuracy and robustness of the model for insulator tiny defects, reduces the error detection rate and missed detection rate, and reduces the calculation amount, and is suitable for small object detection in complex backgrounds.
Smart Images

Figure CN120388007A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention generally relate to the technical field of transmission lines, and particularly to a method and device for detecting insulator defects in transmission lines based on improved Yolo11n. Background Art
[0002] Insulators in transmission lines, as key components in the power system, play an important role in supporting conductors and preventing arc occurrence. However, their performance and reliability are easily affected by factors such as manufacturing defects, environmental erosion, and operation aging. For example, under normal conditions, insulators may accumulate minor damages during long-term operation, the damaged state may be manifested as surface cracks or mechanical damages, and the flashover state may lead to arc discharge due to dirt accumulation or insulation breakdown, thereby triggering system short circuits or large-scale power outages. Therefore, timely and accurate detection and repair of insulator defects are crucial for ensuring the safe operation of the power grid.
[0003] Deep learning has shown significant advantages in the detection of insulator defects in transmission lines. It realizes automatic feature extraction through multi-layer neural networks, without the need for manual feature design, and can efficiently identify complex defect patterns in the four states of normal, damaged, flashover, and self-explosion, with strong adaptability to handle different types and scales of defects.
[0004] Currently, the detection of defects in damaged and flashing insulators in transmission lines still faces many technical bottlenecks, seriously restricting the efficiency and safety level of intelligent operation and maintenance of the power system. First, damaged insulators often appear as tiny cracks or local corners missing in images, with unclear features and easy to blend with complex backgrounds, resulting in low recognition accuracy; while flashing defects such as corona discharge and partial discharge are sudden and short-term, and it is difficult to capture their features based on single-frame images alone, and continuous dynamic information needs to be obtained relying on high-frame-rate videos or special sensors. Second, defect samples are scarce, especially high-quality flashing defect data with labels is difficult to collect, and the cost of manual annotation is high and subjective, restricting the training and generalization ability of deep learning models. At the same time, existing detection models are mostly based on large-scale computing resources, and there are problems of large computational amounts when deployed on unmanned aerial vehicles, inspection robots, or edge devices. Summary of the Invention
[0005] To solve the above problems, in the present invention, the convolutions of the P1, P3, and P5 layers in the backbone feature extraction network module are replaced with DSC_ADown, and different types of features are extracted in a branched manner, enhancing the feature perception ability for tiny targets; the P21 layer in the Neck network is introduced into the ContextGuideFPN context content-guided feature pyramid network module to perform guided fusion of high-level semantic information and low-level detailed features, and combined with the iEMA attention mechanism to guide the interaction and reweighting of features, effectively enhancing the feature expression in the defect area and improving the model's perception ability and detection accuracy for tiny insulator defects in complex backgrounds; the convolutions of the P7 layer in the backbone feature extraction network and the P17 layer in the Neck network are replaced with Dy_SPD to enhance the modeling ability for local details and multi-scale features, further improving the model's stability and accuracy; the loss function CIoU in the original Yolo11n model is replaced with the Wasserstein-IoU loss function to effectively address the detection problem of small target defects in complex backgrounds at a long distance, significantly improving the accuracy and robustness of the model for target localization. The detection accuracy is increased by at least 3% compared with the original Yolo11n model, the detection accuracy is high, and the computational amount of the model is also reduced. Further, the false detection rate and the missed detection rate are reduced.
[0006] According to an embodiment of the present invention, there is provided a method and device for detecting insulator defects on a transmission line based on an improved Yolo11n.
[0007] In a first aspect of the present invention, there is provided a method for detecting insulator defects on a transmission line based on an improved Yolo11n. The method includes:
[0008] Step S01: Collect the original images of the insulator defects on the transmission line and preprocess the images to obtain a data set;
[0009] Step S02: Construct an improved Yolo11n insulator defect detection model for the transmission line: replace the convolutions of the P1, P3, and P5 layers in the Backbone network with the DSC_Adown module; introduce the ContextGuideFPN module into the P21 layer in the Neck network; replace the convolutions of the P7 layer in the Backbone network and the P17 layer in the Neck network with the Dy_SPD module; replace the loss function CIoU in the original Yolo11n model with the Wasserstein-IoU loss function;
[0010] Step S03: Input the data set into the model for training;
[0011] Step S04: Use the trained model to detect the insulator defects on the transmission line to obtain the defect types.
[0012] 2. The method for detecting defects of transmission line insulators based on improved Yolo11n according to claim 1, wherein replacing the convolutional layers of P1, P3, and P5 in the Backbone network with the DSC_Adown module in step S02 is specifically as follows: replacing the 3×3 convolution with a stride of 2 in the P1 layer, P3 layer, and P5 layer with the ADown module, and introducing depthwise separable convolution in the ADown module to form the DSC_ADown module.
[0013] 3. The method for detecting defects of transmission line insulators based on improved Yolo11n according to claim 2, wherein the specific content of the DSC_ADown module is as follows:
[0014] Taking the output of the previous layer as the input, after average pooling, the channels are split into two parts;
[0015] One part performs a 3×3 depth convolution operation with a stride of 2 and a padding of 1, and then a 1×1 pointwise convolution operation with a stride of 1 and a padding of 0. The two are combined to form a depthwise separable convolution;
[0016] The other part first performs a max pooling operation, and then a 1×1 pointwise convolution operation with a stride of 1 and a padding of 0;
[0017] The results of the two parts are concatenated in the channel dimension to generate a new feature map.
[0018] 4. The method for detecting defects of transmission line insulators based on improved Yolo11n according to claim 1, wherein the specific content of the ContextGuideFPN module in step S02 is as follows:
[0019] Receiving input feature maps of different scales, and aligning the number of channels using a 1×1 convolution;
[0020] Concatenating the aligned feature maps in the channel dimension, and sending the concatenated feature maps into the iEMA attention mechanism to automatically learn weights to obtain the attention weights of the original feature maps;
[0021] Implementing channel-level feature selection by multiplying the attention weights and the concatenated feature maps element-wise in the channel dimension to obtain the attention-weighted feature maps;
[0022] Performing bidirectional feature enhancement and information interaction on the attention-weighted feature maps and the original feature maps through residual addition;
[0023] Then performing another channel concatenation and outputting the fused features.
[0024] 5. The method for detecting defects of transmission line insulators based on improved Yolo11n according to claim 1, wherein the specific content of the upper Dy_SPD module in S02 is as follows:
[0025] Automatically compress the spatial dimension of the input feature map using pixel inverse rearrangement, and rearrange the spatial information into the channel dimension;
[0026] The spatial information is globally averaged to obtain a channel-level semantic vector;
[0027] The semantic vector is mapped into an expert weight vector through a fully connected layer and normalized by a Sigmoid activation function to generate a set of dynamic weights in the range of [0, 1];
[0028] Perform 3×3 multi-expert dynamic convolution with a stride of 1 and a padding of 1 on the dynamic weights 4 times to obtain an enhanced feature map with input self-adaptive characteristics;
[0029] Finally, perform normalization and non-linear activation operations to enhance the feature stability and non-linear expression ability.
[0030] 6. The method for detecting defects of transmission line insulators based on improved Yolo11n according to claim 5, wherein the calculation formula of the multi-expert dynamic convolution is as follows:
[0031]
[0032] Where X is the given input feature, Y is the weighted sum of all dynamically selected convolution kernel operations, and W i Is a set of convolution kernels, each convolution kernel corresponds to an expert, and the dynamic coefficient a i Is the dynamic calculation result of the small network, and * represents the convolution operation.
[0033] In the second aspect of the present invention, a device for detecting defects of transmission line insulators based on improved Yolo11n is provided. The device includes:
[0034] Image acquisition module: used to acquire the original image of the transmission line insulator defect and preprocess the image to obtain a data set;
[0035] Model construction module: used to construct an improved Yolo11n transmission line insulator defect detection model: replace the convolutions of P1, P3, and P5 layers in the Backbone network with DSC_Adown modules; introduce the ContextGuideFPN module in the P21 layer of the Neck network; replace the convolutions of the P7 layer in the Backbone network and the P17 layer in the Neck network with Dy_SPD modules; replace the CIoU loss function in the original Yolo11n model with the Wasserstein-IoU loss function;
[0036] Model training module: used to input the dataset into the model for training;
[0037] Defect detection module: used to detect the defects of transmission line insulators using the trained model to obtain the defect types.
[0038] In the third aspect of the present invention, an electronic device is provided. The electronic device includes: a memory and a processor, a computer program is stored on the memory, and when the processor executes the program, the method according to the first aspect of the present invention is implemented.
[0039] In the fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method according to the first aspect of the present invention is implemented.
[0040] In the present invention, by replacing the convolutions of the P1, P3, and P5 layers in the backbone feature extraction network module with DSC_ADown, different types of features are extracted in a branched manner, enhancing the feature perception ability for small targets; introducing the ContextGuideFPN context content-guided feature pyramid network module in the P21 layer of the Neck network, guiding the fusion of high-level semantic information and low-level detailed features, and combining the iEMA attention mechanism to guide the interaction and reweighting between features, effectively enhancing the feature expression in the defect area and improving the model's perception ability and detection accuracy for small insulator defects in complex backgrounds; replacing the convolutions of the P7 layer in the backbone feature extraction network and the P17 layer in the Neck network with Dy_SPD, enhancing the modeling ability for local details and multi-scale features, further improving the model's stability and accuracy; replacing the CIoU loss function in the original Yolo11n model with the Wasserstein-IoU loss function, effectively dealing with the detection problem of small target defects in complex backgrounds at a long distance, and significantly improving the accuracy and robustness of the model for target localization. The detection accuracy is increased by at least 3% compared with the original Yolo11n model, the detection accuracy is high, and the computational amount of the model is also reduced. Furthermore, the false detection rate and the missed detection rate are reduced.
[0041] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key or important features of the embodiments of the present invention, nor to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. Brief Description of the Drawings
[0042] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. Among them:
[0043] Figure 1 Shows a flowchart of a method for detecting defects of transmission line insulators based on improved Yolo11n according to an embodiment of the present invention;
[0044] Figure 2 Shows a schematic diagram of the improved Yolo11n model structure according to an embodiment of the present invention;
[0045] Figure 3 Shows a schematic diagram of the DSC_ADown structure according to an embodiment of the present invention;
[0046] Figure 4 Shows a schematic diagram of the ContextGuideFPN module structure according to an embodiment of the present invention;
[0047] Figure 5 Shows a schematic diagram of the SPD module structure according to an embodiment of the present invention;
[0048] Figure 6 Shows a schematic diagram of the Dy_SPD module structure according to an embodiment of the present invention;
[0049] Figure 7 Shows a detection comparison result diagram according to an embodiment of the present invention;
[0050] Figure 8 Shows a block diagram of a device for detecting defects of transmission line insulators based on improved Yolo11n according to an embodiment of the present invention;
[0051] Figure 9 Shows a schematic diagram of a device for detecting defects of transmission line insulators based on improved Yolo11n according to an embodiment of the present invention. Detailed Description of the Embodiments
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] According to the embodiments of the present invention, a method and device for detecting defects of transmission line insulators based on improved Yolo11n are proposed. By replacing the convolutions of P1, P3, and P5 layers in the backbone feature extraction network module with DSC_ADown, different types of features are extracted in a branching manner, enhancing the feature perception ability for small targets; introducing the P21 layer in the Neck network into the ContextGuideFPN context-guided feature pyramid network module to fuse high-level semantic information and low-level detail features in a guided manner, and combining the iEMA attention mechanism to guide the interaction and reweighting between features, effectively enhancing the feature expression in the defect area and improving the model's perception ability and detection accuracy for small defects of insulators in complex backgrounds; replacing the convolutions of P7 layer in the backbone feature extraction network and P17 layer in the Neck network with Dy_SPD to enhance the modeling ability for local details and multi-scale features, further improving the model's stability and accuracy; replacing the loss function CIoU in the original Yolo11n model with the Wasserstein-IoU loss function to effectively address the detection problem of small target defects in complex backgrounds at long distances, significantly improving the accuracy and robustness of the model for target localization. The detection accuracy is increased by at least 3% compared with the original Yolo11n model, with high detection accuracy, and also reducing the computational amount of the model. Furthermore, the false detection rate and missed detection rate are reduced.
[0054] The principles and spirit of the present invention will be elaborated in detail below with reference to several representative embodiments of the present invention.
[0055] Figure 1 It is a schematic flowchart of a method for detecting defects of transmission line insulators based on improved Yolo11n according to an embodiment of the present invention. The method includes:
[0056] Step S01: Collect the original images of defects of transmission line insulators and preprocess the images to obtain a data set;
[0057] Step S02: Construct an improved Yolo11n transmission line insulator defect detection model: Replace the convolutions of layers P1, P3, and P5 in the Backbone network with the DSC_Adown module; Introduce the ContextGuideFPN module at layer P21 in the Neck network; Replace the convolutions of layer P7 in the Backbone network and layer P17 in the Neck network with the Dy_SPD module; Replace the loss function CIoU in the original Yolo11n model with the Wasserstein-IoU loss function;
[0058] Step S03: Input the dataset into the model for training;
[0059] Step S04: Use the trained model to detect the defects of the transmission line insulators and obtain the defect types.
[0060] It should be noted that although the operations of the method of the present invention are described in a specific order in the above embodiments and accompanying drawings, however, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0061] In order to give a clearer explanation of the above method for detecting defects of transmission line insulators based on the improved Yolo11n, a specific embodiment will be described below. However, it should be noted that this embodiment is only for better explaining the present invention and does not constitute an improper limitation of the present invention.
[0062] The following further illustrates the method for detecting defects of transmission line insulators based on the improved Yolo11n in more detail with a specific example:
[0063] Collect the original images of transmission line insulator defects, process the defect images by means of image rotation, adding noise, weather simulation, color and color gamut enhancement steps to generate images with a size of 600 pixels × 600 pixels, obtain the dataset for training the Yolo11n model, and divide the labeled dataset to obtain the training and validation datasets. Label the transmission line insulator defect images through Labelimg, and the number of labeled defect images is 3840. Divide them into a training set and a validation set according to a ratio of 8:2, with 3072 images in the training set and 768 images in the validation set.
[0064] An improved Yolo11n transmission line insulator defect detection model is constructed: the convolutions of the P1, P3, and P5 layers in the Backbone network are replaced with the DSC_Adown module; the ContextGuideFPN module is introduced in the P21 layer in the Neck network; the convolutions of the P7 layer in the Backbone network and the P17 layer in the Neck network are replaced with the Dy_SPD module; and the CIoU loss function in the original Yolo11n model is replaced with the Wasserstein-IoU loss function.
[0065] like Figure 2 As shown in the figure, in the Backbone feature extraction network, the 3×3 convolutions with a stride of 2 are replaced with ADown modules in the P1, P3, and P5 layers. DepthwiseSeparableConv is introduced into the ADown module to improve it, forming the DSC_ADown module. The DSC_ADown module is designed to reduce the spatial resolution of the feature map while preserving important feature information as much as possible.
[0066] The core idea of the DSC_ADown module is as follows Figure 3 As shown in the figure, the output of the previous layer is used as input. Through channel splitting and parallel processing of two different downsampling paths, AvgPool2d (average pooling) is first performed to achieve preliminary smoothing to reduce local noise interference. The channel is then split into two parts. One part undergoes a 3×3 Depthwise Convolution operation with a stride of 2 and padding of 1, followed by a 1×1 Pointwise Convolution operation with a stride of 1 and padding of 0. The two are combined to form a depthwise separable convolution, which fuses and transforms the channels, independently extracting spatial information from each channel, with lower computational complexity and the ability to expand the receptive field. The other part first undergoes a MaxPool2d (maximum pooling) operation, followed by a 1×1 pointwise convolution operation with a stride of 1 and padding of 0. The results of the two parts are then concatenated channel by channel to generate a new feature map. The DSC_ADown module uses average pooling to preserve global background information and max pooling to highlight salient feature areas. Furthermore, depthwise separable convolution is used to achieve effective channel fusion and compression while reducing computational complexity. This structure effectively enhances the model's ability to perceive defect areas while improving computational efficiency, and minimizes the loss of feature information during downsampling. The number of parameters P for depthwise separable convolution and ordinary convolution is c and computational effort F c The ratio formula is as follows:
[0067]
[0068] Among them, P a , F a are the number of parameters and the amount of computation of depthwise separable convolution, and P b , F b are the number of parameters and the amount of computation of ordinary convolution. The convolution kernel size is D k ×D k , the input channels are C, the output channels are N, and the size of the output feature map is H×W.
[0069] As Figure 2 shown, in the Neck network module, at the P21 layer, the ContextGuideFPN of the context-guided feature pyramid network is used to make up for the lack of full consideration of the semantic alignment and context consistency between feature maps of different scales in the original Yolo11n network, which is likely to cause insufficient fusion of high-level semantics and low-level detail information, especially when dealing with small target defects or complex backgrounds of insulators, there is a problem of accuracy loss. To solve this problem, the ContextGuideFPN module aligns and splices the input scale feature maps, models the global context semantic information through the iEMA channel attention mechanism, and uses the generated attention weights to guide the bidirectional cross-enhancement between the original features. Finally, by splicing the output fusion features, the multi-scale information expression ability in the object detection task is effectively improved, especially suitable for small target and complex background scenarios.
[0070] The core idea of the ContextGuideFPN module is as Figure 4 shown. First, it receives input feature maps of different scales, then uses 1×1 convolution to align the number of channels, splices them in the channel dimension, and sends the spliced feature maps into the iEMA attention mechanism to automatically learn weights, obtaining the attention weights X0_weight and X1_weight of the original feature maps X0 and X1. Through element-wise multiplication of channels, channel-level feature selection is achieved; the channel information with higher weights will be enhanced, and the channel information with lower weights will be suppressed, achieving the purpose of strengthening important features and weakening redundant features. Subsequently, through residual addition, bidirectional feature enhancement and information interaction are realized, that is, using the attention weights of the other party to enhance its own features. Finally, another channel splicing is performed to output more representative fusion features. Compared with traditional direct splicing or addition fusion methods, this module realizes dynamic and efficient information selection and semantic complementarity, and can effectively improve the detection accuracy of small target defects of transmission line insulators.
[0071] As Figure 2As shown in the figure, the 3×3 convolutional downsampling in the P7 layer of the backbone feature extraction network and the P17 layer of the Neck network is replaced by the Dy_SPD module. A downsampling module that combines SPD (Spatial Permutation) technology and DynamicConv (Dynamic Convolution) is proposed. As Figure 5 shown, the original SPD module uses the space-to-Depth (manual spatial permutation) operation to rearrange the spatial dimension of the input to the channel dimension through channel concatenation, and then uses a 3×3 convolution with a stride of 1 and a padding of 1 for feature extraction.
[0072] As Figure 6 shown, the Dy_SPD module automatically compresses the spatial dimensions: height and width of the input feature map by replacing the manual spatial downsampling method with PixelUnshuffle (pixel inverse permutation), and rearranges the spatial information into the channel dimension. After AdaptiveAvgPool (global average pooling), a channel-level semantic vector is obtained, which is then mapped to an expert weight vector through a Linear (fully connected layer) and normalized by a Sigmoid activation function to generate a set of dynamic weights in the range of [0,1]. These weights are used to weighted fuse the convolutional kernels to generate a dynamic convolutional kernel exclusive to the current input sample, thereby realizing an input-aware convolutional operation. The original 3×3 convolution is replaced with 4 3×3 multi-expert dynamic convolutions (CondConv) with a stride of 1 and a padding of 1. The calculation formula of the dynamic convolution is as follows:
[0073]
[0074] where X is the given input feature, Y is the weighted sum of all dynamically selected convolutional kernel operations, W i is a set of convolutional kernels, each convolutional kernel corresponding to an expert, and the dynamic coefficient α i is the dynamic calculation result of a small network, and * represents the convolution operation.
[0075] After the dynamic convolution is completed, the module retains the standard normalization and non-linear activation structures, namely BatchNorm and SiLU activation functions, to enhance the feature stability and non-linear expression ability.
[0076] In the original Yolo11n model, the CIoU loss function is replaced with the Wasserstein-IoU loss function. CIoU is a loss function based on the intersection over union (IoU). In the task of detecting defects in transmission line insulators, there will be targets at a relatively long distance. Since the loss function measured by IoU is very sensitive to the position deviation of small targets, and it may lead to unstable matching and a decline in sample quality in anchor-based detectors, thus affecting the overall detection accuracy. The formula of the CIoU loss function is as follows:
[0077]
[0078] In the formula, IoU is the intersection over union of the predicted box and the ground truth box, c is the diagonal distance of the smallest rectangle that simultaneously covers the target box and the ground truth box, b is the center point of the predicted box, b t is the center point of the ground truth box, ρ is the Euclidean distance calculated between the two center points, (w t ,h t ) are the width and height of the ground truth box, (w, h) are the width and height of the predicted box, α is the weight function, and υ is used to measure the similarity of the aspect ratio.
[0079] To solve this problem, the small target detection evaluation method NWD metric using the Wasserstein distance is used to calculate their similarity through their corresponding Gaussian distributions. Specifically, since the bounding boxes of tiny targets usually show that foreground pixels are dense in the center while background pixels appear more at the edge positions, presenting an obvious characteristic of uneven spatial distribution. To more accurately reflect the importance distribution of each pixel within the bounding box, the bounding box can be modeled as a two-dimensional Gaussian distribution, where the pixels in the central region have the highest weight, and the weight value gradually decays as the distance from the center increases. Then the probability density function of the two-dimensional Gaussian distribution is as follows:
[0080]
[0081] where x, μ, and Σ are the abscissa x, mean vector, and covariance matrix of the Gaussian distribution respectively. x, μ, and Σ satisfy:
[0082] (x - μ) T ∑ -1 (x - μ) = 1
[0083] Therefore, the horizontal bounding box R = (cx, cy, w, h) can be modeled as a two-dimensional Gaussian distribution N(μ, Σ) with the formula:
[0084]
[0085] where, c x and c yDenote the center point coordinates of the horizontal bounding box R; w and h represent the width and height of the horizontal bounding box R respectively.
[0086] In addition, continue to normalize the Gaussian Wasserstein distance. Calculate the annotated bounding box R using the Wasserstein distance a =(cx a , cy a , w a , h a ), and the formula for the Gaussian distribution distance between R b =(cx b , cy b , w b , h b ) is as follows:
[0087]
[0088] where is the distance metric, so it cannot be directly used to measure the similarity between the annotated bounding box R a and R b . Take the exponential to normalize the Wasserstein distance to obtain the formula for the new metric NWD as follows:
[0089]
[0090] where C is a constant, and its value is determined by the dataset.
[0091] Compared with the loss function of the intersection over union, the normalized Wasserstein distance has scale invariance and is smoother for position deviations, and can effectively identify the similarity between non-overlapping or mutually inclusive bounding boxes. Applied to the insulator defect detection task, NWD helps to improve the training efficiency and detection performance of the model. The loss function of NWD is as follows:
[0092] L NWD =1 - NWD(N y , N z )
[0093] where (N y , N z ) are the Gaussian distribution models of the predicted bounding box and the ground truth bounding box respectively.
[0094] As shown in Table 1, the benchmark performance of the unimproved original Yolo11n model on the validation set of transmission line insulator defect detection and the performance of the improved model on the validation set of transmission line insulator defect detection are presented. It can be clearly seen that the precision, recall rate, map0.5, and FLOPS of the improved model have increased by 1.5%, 4.5%, 3.3%, and 1.1% respectively compared to the original model, reducing the number of false detections and missed detections, improving the detection accuracy, and also reducing the computational cost to a certain extent, further enhancing the small target defect detection ability in complex backgrounds.
[0095] Table 1
[0096]
[0097] In summary, the performance has been improved, fully meeting the application requirements of high-precision detection of transmission line insulator images and effectively solving the problem of accurate detection of small target defects in complex backgrounds. The formulas for using the precision P (Precision), recall rate R (Recall), and mean average precision of all classes (map0.5) to evaluate the model performance are as follows:
[0098]
[0099] Among them, TP is the number of correctly predicted positive samples, FN is the number of incorrectly predicted negative samples, FP is the number of incorrectly predicted positive samples, and map is the average detection accuracy of all objects.
[0100] As Figure 7 (a) and Figure 7 (b) show, it can be seen that the improved model can detect small target defects under insulator flashing and breakage that were missed and misdetected by the unimproved model, and the detection accuracy has also been improved to a certain extent in the same environment.
[0101] An insulator defect detection method for transmission lines based on improved Yolo11n of the present invention improves the original Yolo11n model. By replacing the convolutions of P1, P3, and P5 layers in the backbone feature extraction network (Backbone) module with DSC_ADown, different types of features are extracted in a branched manner, enhancing the feature perception ability for small targets and reducing the computational amount, thus reducing the computational amount of the entire module. Introduce the P21 layer in the Neck network into the ContextGuideFPN context-guided feature pyramid network module, fuse the high-level semantic information and low-level detail features in a guided manner, and combine the iEMA attention mechanism to guide the interaction and reweighting between features, effectively enhancing the feature expression in the defect area and improving the model's perception ability and detection accuracy for insulator micro-defects in complex backgrounds. Replace the convolutions of P7 layer in the backbone feature extraction network and P17 layer in the Neck network with Dy_SPD to enhance the modeling ability for local details and multi-scale features, further improving the model stability and accuracy. Replace the loss function CIoU in the original Yolo11n model with the Wasserstein-IoU loss function to effectively handle the detection problem of small target defects in long-distance complex backgrounds, significantly improving the accuracy and robustness of the model for target localization. The detection accuracy is increased by at least 3% compared with the original Yolo11n model, with high detection accuracy. Further, the false detection rate and missed detection rate are reduced.
[0102] Based on the same inventive concept, the present invention also proposes a device for insulator defect detection of transmission lines based on improved Yolo11n. The implementation of this device can refer to the implementation of the above method, and the repeated parts will not be described again. As Figure 2 shown, the device 100 includes:
[0103] An image acquisition module 101: used to acquire the original images of insulator defects on transmission lines and preprocess the images to obtain a data set;
[0104] A model construction module 102: used to construct an improved Yolo11n insulator defect detection model for transmission lines: replace the convolutions of P1, P3, and P5 layers in the Backbone network with DSC_Adown modules; introduce the ContextGuideFPN module into the P21 layer in the Neck network; replace the convolutions of P7 layer in the Backbone network and P17 layer in the Neck network with Dy_SPD modules; replace the loss function CIoU in the original Yolo11n model with the Wasserstein-IoU loss function;
[0105] A model training module 103: used to input the data set into the model for training;
[0106] Defect detection module 104: It is used to detect the defects of transmission line insulators using a trained model and obtain the defect types.
[0107] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0108] As Figure 3 shown, the device includes a central processing unit (CPU), which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.
[0109] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, optical disc, etc.; and a communication unit, such as a network card, modem, wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0110] The processing unit executes the various methods and processes described above, such as method steps S01 to step S04. For example, in some embodiments, method steps S01 to step S04 can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more of the method steps S01 to step S04 described above can be executed. Alternatively, in other embodiments, the CPU can be configured to execute method steps S01 to step S04 in any other appropriate manner (for example, by means of firmware).
[0111] The functions described above herein can be at least partially executed by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.
[0112] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0113] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0114] Furthermore, although the operations are depicted in a particular order, this should be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present invention. Certain features described in the context of separate embodiments can also be implemented in combination in a single implementation. Conversely, the various features described in the context of a single implementation can also be implemented separately or in any suitable sub-combination in multiple implementations.
[0115] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for detecting defects of transmission line insulators based on improved Yolo11n, characterized in that, The method includes: Step S01: Collect the original images of the defects of the transmission line insulators, and preprocess the images to obtain a dataset; Step S02: Construct an improved Yolo11n transmission line insulator defect detection model: Replace the convolutions of the P1, P3, and P5 layers in the Backbone network with the DSC_Adown module; Introduce the ContextGuideFPN module in the P21 layer of the Neck network; Replace the convolutions of the P7 layer in the Backbone network and the P17 layer in the Neck network with the Dy_SPD module; Replace the CIoU loss function in the original Yolo11n model with the Wasserstein-IoU loss function; Step S03: Input the dataset into the model for training; Step S04: Use the trained model to detect the defects of the transmission line insulators to obtain the defect types.
2. The method for detecting defects of transmission line insulators based on improved Yolo11n according to claim 1, characterized in that The replacement of the convolutions of the P1, P3, and P5 layers in the Backbone network with the DSC_Adown module described in Step S02 is specifically as follows: In the P1 layer, P3 layer, and P5 layer, replace the 3×3 convolution with a stride of 2 with the ADown module, and introduce depthwise separable convolutions in the ADown module to form the DSC_ADown module.
3. The method for detecting defects of transmission line insulators based on improved Yolo11n according to claim 2, characterized in that, The specific content of the DSC_ADown module is as follows: Take the output of the previous layer as the input, perform average pooling and then split the channels into two parts; One part performs a 3×3 depth convolution operation with a stride of 2 and a padding of 1, and then a 1×1 pointwise convolution operation with a stride of 1 and a padding of 0. The two are combined to form a depthwise separable convolution; The other part first performs a max pooling operation, and then a 1×1 pointwise convolution operation with a stride of 1 and a padding of 0; Concatenate the results of the two parts in the channel dimension to generate a new feature map.
4. The method for detecting defects of transmission line insulators based on improved Yolo11n according to claim 1, characterized in that, The specific content of the ContextGuideFPN module described in Step S02 is as follows: Receive input feature maps of different scales, and use 1×1 convolutions to align the number of channels; Concatenate the aligned feature maps in the channel dimension, and send the concatenated feature maps into the iEMA attention mechanism to automatically learn the weights to obtain the attention weights of the original feature maps; Perform channel-wise element-wise multiplication on the attention weights and the concatenated feature maps to achieve channel-level feature selection, and obtain the attention-weighted feature maps; Perform bidirectional feature enhancement and information interaction on the attention-weighted feature maps and the original feature maps through residual addition; Perform another channel concatenation and output the fused features.
5. The method for detecting defects of transmission line insulators based on improved Yolo11n according to claim 1, characterized in that The specific content of the upper Dy_SPD module described in S02 is as follows: Automatically compress the spatial dimension of the input feature map using pixel inverse rearrangement, and rearrange the spatial information into the channel dimension; The spatial information passes through global average pooling to obtain a channel-level semantic vector; The semantic vector is mapped to an expert weight vector through a fully connected layer and normalized by the Sigmoid activation function to generate a set of dynamic weights in the range of [0,1]; Perform 4 times of 3×3 multi-expert dynamic convolutions with a stride of 1 and a padding of 1 on the dynamic weights to obtain an enhanced feature map with input self-adaptive characteristics; Finally, normalization and non-linear activation operations are performed to enhance feature stability and non-linear expression ability.
6. The method for detecting defects of transmission line insulators based on improved Yolo11n according to claim 5, characterized in that, The calculation formula of the multi-expert dynamic convolution is as follows: where X is the given input feature, Y is the weighted sum of all dynamically selected convolutional kernel operations, and W i is a set of convolutional kernels, each corresponding to an expert, and the dynamic coefficient a i is the dynamic calculation result of the small network, and * represents the convolutional operation.
7. An apparatus for detecting defects of transmission line insulators based on improved Yolo11n, characterized in that, The device implements the method described in any one of claims 1 to 6, including: Image acquisition module: used to acquire the original images of the defects of transmission line insulators and preprocess the images to obtain a data set; Model construction module: used to construct an improved Yolo11n transmission line insulator defect detection model: replace the convolutions of P1, P3, and P5 layers in the Backbone network with DSC_Adown modules; introduce the ContextGuideFPN module in the P21 layer of the Neck network; replace the convolutions of P7 layer in the Backbone network and P17 layer in the Neck network with Dy_SPD modules; replace the loss function CIoU in the original Yolo11n model with the Wasserstein-IoU loss function; Model training module: used to input the data set into the model for training; Defect detection module: used to detect the defects of transmission line insulators using the trained model to obtain the defect types.
8. An electronic device, comprising a memory and a processor, wherein a computer program is stored on the memory, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1 to 6.
Citation Information
Cited By
Method for detecting defects of stockbridge damper of power transmission line through unmanned aerial vehicle
CN121482644A