Insulator spontaneous explosion defect detection method based on improved YOLOv5

By improving the YOLOv5 detection network model, the lightweight VanillaNet network and depth separation convolution are adopted, and the dynamic sparse attention mechanism and MPDIoU loss function are introduced, which solves the problem of low detection accuracy of insulator self-destruction defects in complex backgrounds, and realizes the lightweight and efficient detection of the model.

CN120198374APending Publication Date: 2025-06-24ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510248820.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the complex context, the detection accuracy of insulator self-destruction defects is low, and the missed detection error rate is high. At the same time, the YOLO series models are large in size, making it difficult to meet the actual deployment and storage needs.

Method used

Improve the YOLOv5 detection network model, adopt the lightweight VanillaNet network as the backbone network, introduce deep separable convolution and partial convolution modules, add dynamic sparse attention mechanism, and use MPDIoU loss function to improve the feature fusion module and detection head.

Benefits of technology

It significantly improves the detection accuracy of self-destruction defects of small-sized insulators under complex backgrounds, reduces the probability of missed detection and false detection, and greatly reduces the parameter quantity and volume of the model, which is suitable for actual deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198374A_ABST
    Figure CN120198374A_ABST
Patent Text Reader

Abstract

The invention discloses an insulator spontaneous explosion defect detection method based on improved YOLOv5, and the method comprises the steps: constructing and training a detection network model of the improved YOLOv5, and replacing a backbone network in a YOLOv5 model with a lightweight VanilaNet network; a weighted feature fusion module is introduced into the neck network, and a fusion path is added for the second scale feature map; a depth separable convolution is adopted to replace an original convolution, a dynamic sparse attention mechanism is introduced, and finally a partial convolution PConv is adopted to replace a last C3 module of the neck network. According to the method, the detection precision of the self-explosion defect of the small-size insulator under the complex background can be effectively improved, the probability of missing detection and false detection is reduced, the parameter quantity and the size of the model are greatly reduced, and the light weight of the model is realized while the detection precision of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of target detection technology, and in particular to a method for detecting insulator self-explosion defects based on improved YOLOv5. Background Art

[0002] In the transmission line system, insulators are key components that bear the important responsibilities of electrical insulation and mechanical connection. However, due to long-term exposure to harsh natural environments, insulators are prone to abnormal problems such as self-explosion, breakage, and flashover. These problems may lead to regional power outages and serious accidents. Therefore, regular inspection of insulators is essential for the normal operation of the power system. There are two main methods for insulator defect detection technology: manual detection and deep learning-based detection.

[0003] Although manual inspection is simple and feasible, it is time-consuming and labor-intensive, and is less flexible due to factors such as the natural environment. In addition, transmission lines are usually distributed in remote areas, making it difficult for manual inspection to cover a wide area. It is inefficient and poses certain safety risks.

[0004] With the development of artificial intelligence and drone technology, target detection models based on deep learning are applied to the operation and maintenance of power systems. By using intelligent devices such as drones embedded with detection models to perform inspection tasks, it can not only ensure personnel safety but also greatly improve detection efficiency.

[0005] The target detection algorithms based on deep learning mainly include one-stage detection algorithms and two-stage detection algorithms. Two-stage detection algorithms (such as R-CNN, Fast R-CNN, Faster R-CNN, etc.) require two-stage processing (candidate area extraction and classification prediction), so the detection accuracy is high, but the structure is complex, the detection speed is slow, and it is difficult to meet the requirements of real-time detection; one-stage detection algorithms (YOLO series, SSD, etc.) directly classify and regress images, and the detection accuracy is relatively low, and it is easy to miss and misdetect, but the detection speed is fast, the network complexity is low, the calculation amount is small, and the generalization and applicability are stronger.

[0006] Among the one-stage detection algorithms, the YOLO series is fast and versatile, and is suitable for insulator defect detection in smart devices such as drones. However, since transmission lines are usually in a complex background and the size of insulator self-explosion defects is small, directly using the YOLO series model for detection is prone to low detection accuracy and missed or false detection of self-explosion defects. In addition, since the storage resources of drones are limited, and the YOLO series model is relatively large, it is easy to cause certain challenges in deployment and storage. Summary of the invention

[0007] The purpose of this application is to provide an insulator self-explosion defect detection method based on improved YOLOv5 to solve the problems of low detection accuracy and high missed detection and false detection rate of small-sized insulator self-explosion defects in complex backgrounds, while realizing the lightweight of the model to solve practical deployment and storage problems.

[0008] To achieve the above purpose, the technical solution of this application is as follows:

[0009] An insulator self-explosion defect detection method based on improved YOLOv5, comprising:

[0010] Build and train the improved YOLOv5 detection network model, including the backbone network, neck network and detection head;

[0011] The image to be detected is passed through the backbone network to extract the first, second, and third scale feature maps, which are then input into the neck network;

[0012] In the neck network, the third-scale feature map is convolved and upsampled, and then weighted feature fused with the second-scale feature map, and then passed through the C3 module and convolution to obtain the second-scale intermediate feature, and then upsampled and weighted feature fused with the first-scale feature map, and then passed through the C3 module to obtain the first-scale output feature; the first-scale output feature is convolved to obtain the first-scale intermediate feature, and the first-scale intermediate feature, the second-scale intermediate feature and the second-scale feature map are weighted feature fused, and then passed through the C3 module to obtain the second-scale output feature; the second-scale output feature is convolved, and then weighted feature fused with the convolved third-scale feature map, and then passed through the feature enhancement module to obtain the third-scale output feature;

[0013] The first, second, and third scale output feature maps are input to their corresponding detection heads to obtain detection results.

[0014] Furthermore, the backbone network is a VanillaNet network.

[0015] Furthermore, the convolution in the neck network is a depth-wise separable convolution.

[0016] Furthermore, the insulator self-explosion defect detection method based on improved YOLOv5 also includes:

[0017] After weighted feature fusion of the first-scale intermediate features, the second-scale intermediate features, and the second-scale feature map, a dynamic sparse attention operation is also performed;

[0018] After the second-scale output features are convolved, they are weighted fused with the convolved third-scale feature map, and a dynamic sparse attention operation is performed.

[0019] Furthermore, the feature enhancement module is a C3 module.

[0020] Furthermore, the feature enhancement module is a partial convolution.

[0021] Furthermore, the training improves the detection network model of YOLOv5 and adopts the MPDIoU loss function.

[0022] This application proposes an insulator self-explosion defect detection method based on improved YOLOv5. Compared with the prior art, the beneficial effects of this application are as follows:

[0023] The backbone network in the YOLOv5 model is replaced by the lightweight VanillaNet network, which discards too many complex operations such as depth, shortcuts, and self-attention, significantly reducing the complexity and volume of the model; the lightweight DWConv module is introduced, and the operation method combining deep convolution and point-by-point convolution is used to further reduce the complexity of the model; the last C3 module of the neck network is replaced by the PConv module. The PConv module utilizes the redundancy in the feature map and only applies conventional convolution to some input channels, and does not operate on the remaining channels, thereby improving computational efficiency and reducing memory access, which can effectively improve the running speed of the model.

[0024] The feature fusion module (Concat) in the original neck network was modified to a weighted feature fusion module (BiFPN_Concat), and a fusion path was added to the second-scale feature map, which effectively improved the model's ability to represent features, its ability to process large-resolution images, and its detection performance of small targets; a dynamic sparse attention mechanism was introduced to enhance the model's detection performance of small targets through dynamic query-aware sparsity; MPDIoU was used to replace the original CIoU in the loss function to solve the problem that the model's original loss function CIoU could not be optimized when the predicted box and the actual annotation box had the same aspect ratio but completely different width and height values, thereby further improving the detection accuracy of the model.

[0025] The improved method proposed in this application can not only effectively improve the detection accuracy of small-sized insulator self-explosion defects under complex backgrounds, but also reduce the probability of missed detection and false detection, and greatly reduce the number of parameters and volume of the model. While improving the detection accuracy of the model, it also achieves the lightweight of the model, which is conducive to the deployment of real application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a flow chart of the insulator self-explosion defect detection method based on improved YOLOv5 in this application.

[0027] Figure 2 A structural diagram of the detection network model for this application.

[0028] Figure 3 This is a schematic diagram of the VanillaNet structure for this application.

[0029] Figure 4 Another structural diagram of the detection network model for this application.

[0030] Figure 5 This is a comparison chart of the evaluation indicators for this application. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0032] Example 1: Figure 1 As shown, a method for detecting insulator self-explosion defects based on improved YOLOv5 is provided, including:

[0033] Step S1, construct and train an improved YOLOv5 detection network model, including a backbone network, a neck network, and a detection head.

[0034] The detection network model constructed in this application is improved on the basis of YOLOv5. YOLOv5 includes a backbone network (Backbone), a neck network (Neck) and a detection head (Head). The neck network is responsible for multi-scale feature fusion, which enhances the feature representation capability by fusing feature maps from different stages of the backbone network; the detection head is responsible for the final target detection.

[0035] It should be noted that YOLOv5 is a relatively mature technology in this field and will not be described in detail here.

[0036] When training the detection network model, this embodiment first selects high-definition images containing insulator self-explosion defects based on the PTL-AI Furnas dataset, removes invalid images, divides them into training sets and validation sets according to a certain ratio, and uses the LabelImg tool to mark the insulator self-explosion defects in the images. Then, the defect dataset is enhanced by using geometric transformation, brightness adjustment, Gaussian blur, and noise addition.

[0037] In a specific embodiment, a total of 1718 images are obtained, 1532 images are selected as a training set, and 186 images are selected as a verification set, covering the categories of insulator self-explosion defects.

[0038] Step S2: The image to be detected is passed through the backbone network to extract the first, second, and third scale feature maps, which are input into the neck network.

[0039] The traditional YOLOv5 object detection network model extracts multi-scale features in its backbone network, usually outputs feature maps of three different scales, and then inputs them into the neck network. The backbone network in this embodiment adopts the backbone network of the original YOLOv5 object detection network model, which will not be elaborated here.

[0040] Step S3: In the neck network, the third-scale feature map is subjected to convolution and upsampling, and then weighted feature fusion is performed with the second-scale feature map. Then, it passes through the C3 module and convolution to obtain the second-scale intermediate feature. After upsampling, weighted feature fusion is performed with the first-scale feature map, and then the first-scale output feature is obtained through the C3 module; the first-scale output feature passes through convolution to obtain the first-scale intermediate feature. The first-scale intermediate feature, the second-scale intermediate feature, and the second-scale feature map are subjected to weighted feature fusion, and then the second-scale output feature is obtained through the C3 module; the second-scale output feature passes through convolution and then weighted feature fusion with the convolved third-scale feature map, and then the third-scale output feature is obtained through the feature enhancement module.

[0041] This embodiment improves the neck network in the YOLOv5 object detection network model. Specifically, the feature fusion module (Concat) in the original neck network is modified to a weighted feature fusion module (BiFPN_Concat), and a fusion path is added for the second-scale feature map, enabling the network to fuse more features without adding too much cost.

[0042] In the neck network, the third-scale feature map is subjected to convolution and upsampling, and then weighted feature fusion is performed with the second-scale feature map. Then, it passes through the C3 module and convolution to obtain the second-scale intermediate feature. After upsampling, weighted feature fusion is performed with the first-scale feature map, and then the first-scale output feature is obtained through the C3 module; the first-scale output feature passes through convolution to obtain the first-scale intermediate feature. The first-scale intermediate feature, the second-scale intermediate feature, and the second-scale feature map are subjected to weighted feature fusion, and then the second-scale output feature is obtained through the C3 module; the second-scale output feature passes through convolution and then weighted feature fusion with the convolved third-scale feature map, and then the third-scale output feature is obtained through the feature enhancement module.

[0043] Specifically, as Figure 2As shown: A new fusion path is added between the third layer and the fourteenth layer, which can enable the network to fuse more features without adding too much cost; the Concat module located at the seventh layer is replaced with the BiFPN_Concat module, and the BiFPN_Concat module performs weighted feature fusion on the output feature map of the third layer and the output feature map of the sixth layer; the Concat module located at the eleventh layer is replaced with the BiFPN_Concat module, and the BiFPN_Concat module performs weighted feature fusion on the output feature map of the second layer and the output feature map of the tenth layer; the Concat module located at the fourteenth layer is replaced with the BiFPN_Concat module, and the BiFPN_Concat module performs weighted feature fusion on the output feature map of the third layer, the output feature map of the ninth layer, and the output feature map of the thirteenth layer; the Concat module located at the eighteenth layer is replaced with the BiFPN_Concat module, and the BiFPN_Concat module performs weighted feature fusion on the output feature map of the fifth layer and the output feature map of the seventeenth layer.

[0044] Among them, the seventh layer, the eleventh layer, the fourteenth layer, and the eighteenth layer are weighted feature fusion modules (BiFPN_Concat); the third-scale feature map passes through the convolution of the fifth layer and the upsampling of the sixth layer, and is input to the seventh layer to perform weighted feature fusion with the second-scale feature map; the fused features pass through the C3 module of the eighth layer and the convolution of the ninth layer to output the second-scale intermediate features; then through the upsampling of the tenth layer, input to the eleventh layer, and perform weighted feature fusion with the first-scale feature map; then through the C3 module of the twelfth layer to output the first-scale output features; then through the convolutional layer of the thirteenth layer to output the first-scale intermediate features; the first-scale intermediate features, the second-scale intermediate features, and the second-scale feature map perform weighted feature fusion at the fourteenth layer; then through the C3 module of the sixteenth layer to output the second-scale output features; after convolution through the seventeenth layer and input to the eighteenth layer, at the same time the third-scale feature map also passes through the convolution of the fifth layer and is input to the eighteenth layer, and after weighted feature fusion in the eighteenth layer, through the C3 module of the twentieth layer to obtain the third-scale output features.

[0045] The weighted feature fusion module adopts the weighted feature fusion method of fast normalization fusion, which can assign a learnable weight to each fusion path, enabling the network to learn the importance of different input features through the weights and fuse different input features in a differentiated manner. This method scales the weights to between [0,1], with fast training speed and high efficiency, and its expression is shown in Equation (1):

[0046]

[0047] Among them, w irepresents the learnable weight, which is always in the interval (0,1); w j is the learning weight of the jth layer; I i is the i-th input feature; ∈ is the minimum parameter, which takes a value of 0.0001.

[0048] This embodiment effectively improves the model's feature fusion capability through weighted feature fusion, new fusion paths, etc., enriches the information contained in the feature graph, and thus improves the model's detection capability for small-sized insulator self-explosion defects.

[0049] It should be noted that in the original YOLOv5, the feature enhancement module of the last layer is the C3 module.

[0050] Step S4: The first, second and third scale output feature maps are input to their corresponding detection heads to obtain detection results.

[0051] The detection head of this embodiment still uses the original YOLOv5 detection head. The detection head processes the output features of each scale to obtain the detection result, which will not be repeated here.

[0052] Example 2: This example further improves on Example 1, specifically by replacing the Backbone network in the YOLOv5 model with a lightweight VanillaNet network, simplifying the network structure of the model to reduce the size and complexity of the model.

[0053] That is, the backbone network described in this embodiment is a VanillaNet network.

[0054] exist Figure 2 The Backbone network in the YOLOv5 model has been replaced with a lightweight VanillaNet network. Figure 3As shown in the figure, the VanillaNet network can be divided into four stages. In stage 1, a 4×4 convolutional layer with a stride of 4 is used to perform feature mapping on the three-channel image and downsample it, reducing the spatial dimension of the image while increasing the number of channels. In stage 2, 3 1×1 convolutional layers are used to maintain the feature map information while reducing the computational cost. There is a batch normalization layer after each convolutional layer to accelerate the training process and stabilize the training. And a max pooling layer with a stride of 2 is used for feature downsampling, reducing the spatial dimension of the feature map and increasing the number of channels. An activation function is applied after each convolutional layer to enhance the nonlinear ability of the network. In stage 3, 1 1×1 convolutional layer is used to keep the number of channels unchanged, and an average pooling layer is used to further reduce the spatial dimension of the feature map. In stage 4, a fully connected layer is used to map the high-dimensional features to specific classification labels and output the classification results. The VanillaNet network abandons excessive complex operations such as excessive depth, shortcut, and self-attention, and uses an extremely simple network structure, which has a lower number of parameters and complexity compared to the Backbone backbone network in the original model, and can effectively reduce the volume and complexity of the model.

[0055] As Figure 2 shown, the second, third, and fourth stages, namely the second, third, and fourth layers, respectively output the first, second, and third scale feature maps.

[0056] Example 3: This example makes further improvements on the basis of Examples 1 and 2. Specifically, the original convolutional Conv module in the Neck network of the YOLOv5 model is replaced with a depthwise separable convolutional DWConv module, reducing the computational amount and the number of parameters of the model, while improving the efficiency and speed of the model.

[0057] That is, the convolution in the neck network is a depthwise separable convolution.

[0058] During the convolution operation in the Conv module, the convolution kernel extracts the features of all channels and fuses them. The input is a feature map of size a×b×c, the Filter is p, the convolution kernel size is q×q×c, and the output size is a×b×p feature map. The expressions for the number of parameters ConvP and the computational amount ConvC required during the convolution operation are shown in Equations (2) and (4) respectively:

[0059] ConvP = q×q×c×p (2)

[0060] ConvC = a×b×q×q×c×p (3)

[0061] The DWConv module decomposes the traditional convolution operation process into two steps, namely, depth convolution and point-by-point convolution. The input feature map of size a×b×c is firstly performed with depth convolution and convolution kernel size q×q×c. Each channel of the input layer is independently convolved to generate a feature map. Then, point-by-point convolution is performed with convolution kernel size 1×1×c and number p. The feature map generated by the depth convolution operation is weightedly combined in the depth direction to generate a new feature map of size a×b×p. The expressions of the parameter amount DWConvP and the calculation amount DWConvC required in the convolution operation are shown in the following equations (4) and (5), respectively:

[0062] DWConvP=q×q×c+c×p (4)

[0063] DWConvC=a×b×q×q×c+a×b×c×p (5)

[0064] It can be seen from the above formula that compared with the Conv module, the convolution operation of the DWConv module has fewer calculation parameters and a smaller amount of calculation. Using the DWConv module can not only reduce the amount of calculation and parameters of the model, but also improve the efficiency and speed of the model.

[0065] Example 4: This example further improves the neck network by introducing a dynamic sparse attention mechanism (BiLevelRoutingAttention) into the neck network, utilizing its dynamic query-aware sparsity to improve the flexibility of the model in terms of computational allocation and content perception, and enhance the model's detection performance for small targets.

[0066] That is, a method for detecting insulator self-explosion defects based on improved YOLOv5, further comprising:

[0067] After weighted feature fusion of the first-scale intermediate features, the second-scale intermediate features, and the second-scale feature map, a dynamic sparse attention operation is also performed;

[0068] After the second-scale output features are convolved, they are weighted fused with the convolved third-scale feature map, and a dynamic sparse attention operation is performed.

[0069] Specifically, Figure 2 As shown, the fifteenth layer is introduced between the fourteenth and sixteenth layers, and the nineteenth layer is introduced between the eighteenth and twentieth layers, corresponding to Figure 2 The BiLevelRoutingAttention module in the BiLevelRoutingAttention module performs dynamic sparse attention operations and has a structure like Figure 4 As shown, it is a dynamic, query-aware sparse attention mechanism.

[0070] Example 5: This example further improves the neck network. As Figure 4 shown, use the partial convolution PConv to replace the last C3 module of the neck network, reducing the computational redundancy and memory access of the model and improving the running speed of the model.

[0071] The partial convolution PConv processes through the identity mapping and the convolutional layer in parallel, and then adds the outputs of these two paths to form the final output. By exploiting the redundancy in the feature map and only applying the conventional convolution to some of the input channels without operating on the remaining channels, the PConv module reduces the redundant computation and memory access, has a smaller FLOPs and memory access volume, and can effectively improve the running speed of the model.

[0072] It should be noted that Figure 4 shows the detection network model including all improvements. This application does not list all possible combinations of improvement embodiments. Those skilled in the art can combine various improvements for implementation, which will not be elaborated here.

[0073] Example 6: When training the detection network model of this application, replace the loss function CIoU of the original model with MPDIoU to improve the convergence speed and detection accuracy of the model. The calculation process of the MPDIoU loss is shown in the following formulas (6) to (9):

[0074] d1 2 =(x1 B -x1 A ) 2 +(y1 B -y1 A ) 2 (6)

[0075] d2 2 =(x2 B -x2 A ) 2 +(y2 B -y2 A ) 2 (7)

[0076]

[0077] MPDIoU Loss = 1 - MPDIoU (9)

[0078] where A is the ground truth box, B is the predicted box, (x1 A ,y1 A ),(x2 A ,y2 A ) represent the coordinates of the upper left and lower right corner points of the ground truth box A, (x1 B ,y1B ),(x2 B ,y2 B ) represents the coordinates of the upper left and lower right corners of the predicted box B, w and h are the width and height of the input image, d1 and d2 represent the distance between the upper left and lower right corners of the real box A and the predicted box B. The loss function MPDIoU is used to minimize the distance between the upper left and lower right corners of the predicted bounding box and the actual annotated bounding box, solving the problem that the original loss function CIoU of the model cannot be optimized when the predicted box and the actual annotated box have the same aspect ratio but completely different width and height values, thereby improving the detection performance of the model.

[0079] This application conducted a comparative experiment on the original YOLOv5 model and the improved detection network model of this application, setting the initial learning rate and final learning rate to 0.01, the momentum to 0.937, the optimizer to SGD, the image size to 640×640, the batch size to 64, and the number of training rounds to 200. The comparison indicators selected were evaluation indicators such as precision, recall, mean average precision (mAP), and model parameters (Params). The experimental results are shown in Table 1:

[0080] Table 1

[0081]

[0082] As shown in Table 1, the values ​​of the key evaluation indicators of the improved YOLOv5 detection network model proposed in this application are 92.0%, 84.5%, 90.3%, 47.7% and 3.4M respectively in Precision, Recall, mAP50, mAP50:95, Params, etc. Compared with the original model, the Recall index is improved by 3.7%, which can effectively reduce the phenomenon of missed detection and false detection of small-sized insulator self-explosion defects under complex backgrounds. In addition, the overall parameter volume is reduced from 7.0M to 3.4M, a decrease of 51.4%, and the volume is reduced from 14.3MB to 7.0MB, with lower complexity and smaller volume, and greater advantages in equipment deployment and storage. Although the Precision and mAP50:95 indicators decreased by 2.9% and 0.4% respectively, the mAP50 indicator increased by 1%, indicating that the detection accuracy of the improved YOLOv5 detection network model has been improved, and small-sized insulator self-explosion defects under complex backgrounds can be better identified. This application detects the network model and YOLOv5s training results. Figure 5 As shown, although the Precision and mAP50:95 curves have decreased, the Recall and mAP50 curves have increased, and the overall detection capability of the model has improved. This shows that the improved method proposed in this application has improved the detection accuracy of the model while achieving model lightweight.

[0083] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A method for detecting insulator self-explosion defects based on improved YOLOv5, characterized in that: The insulator self-explosion defect detection method based on improved YOLOv5 includes: Build and train the improved YOLOv5 detection network model, including the backbone network, neck network and detection head; The image to be detected is passed through the backbone network to extract the first, second, and third scale feature maps, which are then input into the neck network; In the neck network, the third-scale feature map is convolved and upsampled, and then weighted feature fused with the second-scale feature map, and then passed through the C3 module and convolution to obtain the second-scale intermediate feature, and then upsampled and weighted feature fused with the first-scale feature map, and then passed through the C3 module to obtain the first-scale output feature; the first-scale output feature is convolved to obtain the first-scale intermediate feature, and the first-scale intermediate feature, the second-scale intermediate feature and the second-scale feature map are weighted feature fused, and then passed through the C3 module to obtain the second-scale output feature; the second-scale output feature is convolved, and then weighted feature fused with the convolved third-scale feature map, and then passed through the feature enhancement module to obtain the third-scale output feature; The first, second, and third scale output feature maps are input to their corresponding detection heads to obtain detection results.

2. The insulator self-explosion defect detection method based on improved YOLOv5 according to claim 1 is characterized in that: The backbone network is the VanillaNet network.

3. The insulator self-explosion defect detection method based on improved YOLOv5 according to claim 1 or 2, characterized in that: The convolution in the neck network is a depth-wise separable convolution.

4. The insulator self-explosion defect detection method based on improved YOLOv5 according to claim 3 is characterized in that: The insulator self-explosion defect detection method based on improved YOLOv5 also includes: After weighted feature fusion of the first-scale intermediate features, the second-scale intermediate features, and the second-scale feature map, a dynamic sparse attention operation is also performed; After the second-scale output features are convolved, they are weighted fused with the convolved third-scale feature map, and a dynamic sparse attention operation is performed.

5. The insulator self-explosion defect detection method based on improved YOLOv5 according to claim 1 is characterized in that: The feature enhancement module is a C3 module.

6. The insulator self-explosion defect detection method based on improved YOLOv5 according to claim 1 is characterized in that: The feature enhancement module is a partial convolution.

7. The insulator self-explosion defect detection method based on improved YOLOv5 according to claim 1 is characterized in that: The training improves the detection network model of YOLOv5 and adopts the MPDIoU loss function.