Power transmission line foreign matter detection method based on multi-scale fusion enhanced self-adaption

By building the YOLOv8-ARD model, the problems of low manual detection efficiency and inaccurate small target detection in transmission line foreign object detection are solved, achieving more efficient and accurate foreign object detection to meet the needs of safe and stable operation of the power grid.

CN120708020APending Publication Date: 2025-09-26BAISHAN POWER SUPPLY COMPANY OF STATE GRID JILIN ELECTRONICS POWER COMPANY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510657488.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing methods for detecting foreign objects on transmission lines rely on manual inspection, which is inefficient, costly, and susceptible to human error. Drone-based detection is inaccurate in detecting small targets in complex backgrounds, has high computational complexity, and is difficult to meet the needs of safe and stable operation of the power grid.

Method used

The YOLOv8-ARD model is constructed based on the YOLOv8 model. By replacing the downsampling layers on the Backbone and Head ends with ADown modules, replacing the C2F module with the RepNCSPELAN4 module, and adding an auxiliary detection head DetectAux on the Head end, the loss function is optimized to improve detection accuracy and reduce computational complexity.

Benefits of technology

The accuracy and regression rate of foreign object detection on transmission lines are improved, the computational complexity is reduced, and the model's ability to detect small targets is enhanced. It has good robustness and generalization and is suitable for complex field environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708020A_ABST
    Figure CN120708020A_ABST
Patent Text Reader

Abstract

The invention discloses a power transmission line foreign matter detection method based on multi-scale fusion enhanced self-adaption, belongs to the field of power systems, and aims to improve the accuracy of power transmission line foreign matter intrusion detection. The method comprises the following detection steps: obtaining an image data set of foreign matter invasion of the power transmission line, and screening and classifying; manually marking the images and dividing the images into a training set and a verification set; a YOLOv8-ARD model is constructed based on the YOLOv8 model, and a lower sampling layer of a Backbone end and a lower sampling layer of a Head end of the YOLOv8 model are replaced by an ADown module so as to fuse multi-scale features; a C2F module in the YOLOv8 model is replaced by a RepNCSPELAN4 module, so that the feature extraction efficiency and the reasoning speed are improved; an auxiliary detection head is added to the detection head part of the Head end to assist the main detection head in completing a small target detection task; training a YOLOv8-ARD model: inputting a training set in the marked data set into the YOLOv8-ARD model for training, and obtaining a YOLOv8-ARD.pt weight file; and using the trained YOLOv8-ARD model to carry out fault detection and positioning on the foreign matters in the foreign matter invasion picture of the power transmission line to be identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power systems, and in particular relates to a transmission line foreign object detection method based on multi-scale fusion enhanced adaptation. Background Art

[0002] Overhead transmission lines are crucial channels for the transmission of electric energy and are directly linked to the safe and stable operation of the power grid. Due to the complex environments through which transmission lines pass, foreign object intrusion is a major cause of line tripping and failure to reclose successfully. Foreign object intrusion, such as kites, debris, and mechanical construction, can cause power line short circuits, shorten discharge distances, and thus impact the stable operation of the power system. Traditional defect detection methods rely primarily on manual inspection. While widely adopted, this approach has several significant limitations, including low efficiency, high labor costs, and susceptibility to human error. Furthermore, the remote and often complex terrain in which transmission lines operate further complicates the inspection process. In recent years, the emergence of drone-based inspection systems has revolutionized transmission line maintenance. Drones provide unparalleled access to hard-to-reach areas and can capture high-resolution images, significantly improving inspection efficiency and coverage. However, detecting defects in drone-captured images presents significant challenges due to complex backgrounds, densely populated areas, small objects, and occlusions.

[0003] To address these challenges, deep learning-based object detection algorithms have gained significant attention due to their superior performance in image analysis tasks. Unlike traditional image processing methods, deep learning algorithms, especially convolutional neural networks (CNNs), excel in automatically learning discriminative features of large datasets, thereby achieving significant progress in detection accuracy and efficiency. However, CNNs typically contain a large number of convolutional layers and parameters, which results in very high computational complexity and memory requirements. For example, a typical CNN model may contain millions or even tens of millions of parameters, which makes the training and inference processes very time-consuming. SSD (Single Shot MultiBox Detector) is a single-stage detection algorithm that performs detection on feature maps of different scales. However, since the SSD algorithm uses anchor points on the feature map to detect targets, the detection of small targets is not accurate enough. Summary of the Invention

[0004] In view of this, the present invention provides a transmission line foreign object detection method based on multi-scale fusion enhanced adaptation, which constructs the YOLOv8-ARD model based on the YOLOv8 model to improve the detection accuracy of foreign object intrusion in the transmission line, improve the regression rate of detection and reduce the amount of calculation, so as to more accurately detect and locate foreign objects in the transmission line.

[0005] The technical solution adopted by the present invention to solve the above technical problems is:

[0006] A method for detecting foreign objects on power transmission lines based on multi-scale fusion and enhanced self-adaptation includes the following detection steps:

[0007] S1, obtain the image dataset of foreign objects intruding on the transmission line and screen and classify it;

[0008] S2, manually annotate the screened foreign body invasion images and divide them into training set and validation set;

[0009] S3: Build the YOLOv8-ARD model based on the original YOLOv8 model. The original YOLOv8 model includes the Backbone end for feature extraction, the Neck end for feature fusion, and the Head end for prediction. The process of building the YOLOv8-ARD model includes:

[0010] The downsampling layers on the Backbone and Head sides of the YOLOv8 model are replaced with ADown modules. The C2F module in the YOLOv8 model is replaced with the RepNCSPELAN4 module. An auxiliary detection head is added to the detection head on the Head side to assist the main detection head in completing the detection task of small objects.

[0011] S4, training YOLOv8-ARD model: input the training set in the labeled dataset into the YOLOv8-ARD model for training to obtain the YOLOv8-ARD.pt weight file;

[0012] S5 uses the trained YOLOv8-ARD model to detect and locate foreign objects in the transmission line foreign object intrusion image to be identified.

[0013] Furthermore, the implementation process of step S1 is as follows:

[0014] S11, using unmanned inspection equipment to collect images of foreign objects intruding on the transmission line;

[0015] S12, manually screening the acquired images and obtaining an image data set, wherein during the manual screening, intact images of foreign objects intruding the transmission line are removed, and images of foreign objects intruding the transmission line with obstructions are retained;

[0016] S13, classifying the screened transmission line foreign body images.

[0017] Furthermore, in S2, the process of manually labeling the screened foreign body intrusion images and dividing them into a training set and a validation set is as follows:

[0018] S21, using an image annotation tool and selecting the YOLO annotation format to annotate the classified foreign body intrusion image;

[0019] S22, save each label name and its corresponding bounding box into a .txt file, the saved file forms a dataset, and use a Python script to randomly divide the dataset into a training set and a validation set in a ratio of 8:2.

[0020] Furthermore, in S3, the ADown module includes two branches, namely cv1 and cv2. cv1 uses a 3x3 convolution kernel for downsampling, and cv2 first performs 3x3 maximum pooling and then uses a 1x1 convolution kernel for downsampling. The ADown module processes the data set as follows: in the forward propagation, the feature map is input into the ADown module, the feature map is downsampled by average pooling, and the downsampled feature map is divided into two parts, one part is processed by cv1 and the other part is processed by cv2, and finally the two parts are spliced ​​together.

[0021] Furthermore, in the S3, the parameters of the RepNCSPELAN4 module are set to C1, C2, C3, C4 and C5, where C1 is the number of input channels, C2 is the number of output channels, C3 is the number of intermediate feature channels, usually half of C2, C4 is the number of channels in the RepNCSP module, and C5 is the number of times the RepNCSP module is repeated, which defaults to 1; the structure of the RepNCSPELAN4 module includes self.cv1, self.cv2, self.cv3 and self.cv4, self.cv1 input The input first passes through a 1x1 convolution, and the input features are converted from C1 channels to C3 channels; then the feature map is divided into two parts; self.cv2 contains a RepNCSP module and a 3x1 convolution module. The RepNCSP module converts the features of the C3 / 2 channels to the C4 channels, and then keeps the number of channels unchanged through 3x1 convolution; self.cv3 contains a RepNCSP module and a 3x1 convolution module; self.cv4 includes a 1x1 convolution, which converts the merged features from C3+2*C4 channels to C2 channels.

[0022] Furthermore, in S3, the Head end uses three DetectAux modules as auxiliary detection heads.

[0023] Furthermore, during the training of the YOLOv8-ARD model, the loss function of the auxiliary detection head DetectAux includes classification loss and bounding box regression loss;

[0024] The loss function is expressed as:

[0025] Loss=BCEcls(d1)+BCEcls(d2)+CIoU(d1)+CIoU(d2)+DFL(d1)+DFL(d2)

[0026] Among them, BCEcls represents binary cross entropy loss, which is used for classification tasks; CIoU represents CIoU loss, which is used for bounding box regression tasks; DFL represents distribution focus loss, which is used for bounding box regression tasks; d1 and d2 represent the outputs of the main detection head and auxiliary detection head, respectively;

[0027] Classification loss is a function used to measure the difference between the classification model prediction results and the true label. The formula is expressed as:

[0028]

[0029] Among them, y i is the true label of the i-th sample, is the predicted label of the i-th sample, and N is the number of samples.

[0030] The formula for bounding box regression loss is expressed as:

[0031]

[0032] Among them, IoU(b,b * ) is the predicted box b and the real box b * The intersection-and-union ratio of p 2 (b,b * ) is the square of the Euclidean distance between the center point of the predicted box and the real box; c is the diagonal length of the minimum closure containing the predicted box and the real box; ar(b) and ar(b * ) are the aspect ratios of the predicted box and the true box respectively; α is an adjustment parameter used to balance the weights of different parts;

[0033] Distribution focus loss is used for bounding box regression tasks. The formula of distribution focus loss is:

[0034]

[0035] Where y is the true label, y^ is the predicted label, and N is the number of samples.

[0036] The beneficial effects of the present invention compared with the prior art are:

[0037] The YOLOv8-ARD model constructed by the present invention is obtained by replacing the downsampling layers of the Backbone and Head ends of the YOLOv8 model with ADown modules, replacing the C2F module in the YOLOv8 model with the RepNCSPELAN4 module, and adding an auxiliary detection head DetectAux to the detection head part of the Head end. The ADown module can gradually extract features from the feature map to fuse multi-scale features. The RepNCSPELAN4 module can enhance the model's ability to recognize targets of different sizes. By adding auxiliary detection heads, the detection process of the model can be decomposed into multiple subtasks, and each detection head can be optimized and adjusted independently. The resulting YOLOv8-ARD model has stronger feature extraction capabilities, better performance in distinguishing similar features, more thorough multi-scale feature fusion, stronger ability to detect small targets, and lower missed detection rate and false detection rate. Even in complex natural environments in the wild, it has good detection effects, good robustness and generalization, and can meet the requirements of unmanned inspections. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings are incorporated in and constitute a part of this application and are used to provide a further understanding of the present invention.

[0039] Figure 1 This is a flow chart of the foreign object intrusion detection process for power transmission lines according to this embodiment.

[0040] Figure 2 Schematic diagram of the structure of the YOLOv8-ARD model of this embodiment.

[0041] Figure 3 A schematic diagram of using the labelimg image annotation tool to annotate the dataset.

[0042] Figure 4 Schematic diagram of the structure of the RepNCSPELAN4 module.

[0043] Figure 5 Schematic diagram of foreign object intrusion detection for power transmission lines according to this embodiment.

[0044] Figure 6 This is a comparison chart of the detection results of the YOLOv8 model and the YOLOv8-ARD model for insulator equipment detection in the presence of artificial high-altitude foreign objects on transmission lines.

[0045] Figure 7 This is a comparison chart of the detection results of the YOLOv8-ARD model in this embodiment and other YOLO models. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. The following embodiments are used to illustrate the present invention but are not used to limit the scope of the present invention.

[0047] Figure 1 The flow chart of the foreign body intrusion detection for the power transmission line of this embodiment is shown as follows: Figure 1 As shown, the present embodiment provides a method for detecting foreign objects on a power transmission line based on multi-scale fusion enhanced adaptation, comprising the following detection steps:

[0048] S1, obtain the image dataset of foreign objects intruding on the transmission line and screen and classify them:

[0049] S11. Use unmanned inspection equipment, such as drones, to collect images (visible light images) of foreign objects intruding on transmission lines in a field environment. When taking images of foreign objects intruding on transmission lines, sufficient lighting conditions are required to ensure that the foreign objects on the transmission lines being photographed are avoided as much as possible.

[0050] Foreign objects captured on power transmission lines can include animals, high-altitude intruders, artifacts of human activity, and garbage. Birds and other animals nesting on transmission lines or towers can affect the safe operation of the lines. High-altitude intruders such as balloons can cause line failures and even local or regional power outages. Human artifacts such as kites can short-circuit power lines, shortening discharge distances and affecting the stable operation of the power system. In severe cases, they can even cause tripping and paralysis. Plastics, paper scraps, metal fragments, and other garbage can be blown onto or become suspended on transmission lines due to wind, human activity, or other reasons, posing a threat to the safe and stable operation of the power system. Therefore, these foreign objects are captured as subjects.

[0051] S12, manually screening the acquired images and obtaining an image data set, wherein the manual screening method is to remove intact transmission line foreign body intrusion images and retain obstructed (damaged) transmission line foreign body intrusion images.

[0052] S13, classifying the screened transmission line foreign body images, classifying animals such as birds into one category, high-altitude intruders such as balloons or hot air balloons into another category, products of human activities such as kites into another category, and garbage and waste such as plastic products, paper scraps, and metal fragments into another category.

[0053] S2: Manually label the selected foreign body intrusion images and divide them into training set and validation set:

[0054] S21, use image annotation tools, such as labelimg software, and select YOLO annotation format to annotate the classified foreign body invasion images. The specific method of data annotation is as shown in the attached Figure 3 As shown, Figure 3 For the image of foreign objects intruding into a damaged power transmission line, the entire image of foreign objects intruding into the power transmission line is labeled. The label name of the animal is set to nest, the label name of the high-altitude intrusion object is set to balloon, the label name of the product of human activity is set to kite, and the label name of the garbage and waste is set to trash.

[0055] S22, save each label name and its corresponding bounding box into a .txt file, the saved file forms a dataset, and use a Python script to randomly divide the dataset into a training set and a validation set in a ratio of 8:2.

[0056] S3, constructing a YOLOv8-ARD model to perform fault detection and positioning on the transmission line foreign body intrusion image to be identified; although the backbone network of the YOLOv8 original model has a strong feature extraction capability, in the wild natural environment, the aerial images of transmission line foreign body intrusion are easily affected by factors such as lighting, occlusion, complex background, and the small area of ​​the transmission line foreign body intrusion, which makes the collected images difficult to identify; Therefore, this embodiment constructs a YOLOv8-ARD model based on the YOLOv8 original model to improve the detection accuracy of transmission line foreign body intrusion, improve the regression rate of detection and reduce the amount of calculation, so as to more accurately detect and locate transmission line foreign bodies. Among them, the YOLOv8 model includes a Backbone end for feature extraction, a Neck end for feature fusion, and a Head end for prediction. The specific construction process of the YOLOv8-ARD model is as follows:

[0057] S31, replace the downsampling layers of the Backbone and Head ends of the YOLOv8 model with the ADown module: Figure 2 The location of ADown module replacement is marked in detail, such as Figure 2As shown in the figure, on the Backbone side, starting with the input image, after passing through the first CBS module, features are extracted and the number of channels is adjusted, ultimately reaching the SPPF module to fuse multi-scale features. The ADown module consists of two main branches, cv1 and cv2. cv1 uses a 3x3 convolution kernel for downsampling, while cv2 performs a 3x3 max pooling followed by a 1x1 convolution kernel for downsampling. In the forward propagation, the feature map is input to the ADown module, which downsamples it using average pooling. The downsampled feature map is then split into two parts, one processed by cv1 and the other by cv2. Finally, the two results are concatenated. The ADown module reduces the complexity of the YOLOv8 model by reducing the number of parameters, which helps improve its efficiency, especially in resource-constrained environments. Although ADown aims to reduce the spatial resolution of the feature map, its design also focuses on preserving as much image information as possible, enabling the YOLOv8 model to more accurately detect objects.

[0058] S32, replace the C2F module in the YOLOv8 model with the RepNCSPELAN4 module to improve feature extraction efficiency and inference speed: Figure 2 As shown, Figure 2 The replacement location for the RepNCSPELAN4 module is clearly marked. In actual transmission line inspection scenarios, poor insulator image quality can occur due to issues like shooting angle, shadows, and insufficient lighting. For such images, the YOLOv8 model cannot extract meaningful features, resulting in poor feature fusion and even impacting the model's learning ability. The RepNCSPELAN4 module is a composite feature extraction and fusion module that offers advantages over the C2F module in terms of non-local connections, cross-scale exchange, parameter reparameterization, and the ELAN4 architecture. Therefore, replacing the C2F module with the RepNCSPELAN4 module makes the YOLOv8 model more efficient and powerful in feature extraction and fusion. The parameters of the RepNCSPELAN4 module are C1, C2, C3, C4, and C5. C1 is the number of input channels, C2 is the number of output channels, C3 is the number of intermediate feature channels (typically half of C2), C4 is the number of channels in the RepNCSP module, and C5 is the number of RepNCSP module repetitions, which generally defaults to 1.

[0059] Figure 4 The schematic diagram of the structure of the RepNCSPELAN4 module is shown in FIG. Figure 4 As shown, the structure of the module is:

[0060] self.cv1: The input first undergoes a 1x1 convolution, and the input features are converted from C1 channels to C3 channels; then the feature map is divided into two parts.

[0061] self.cv2: Contains a RepNCSP module and a 3x1 convolution module. The RepNCSP module converts the features of the C3 / 2 channels into C4 channels, and then keeps the number of channels unchanged through 3x1 convolution.

[0062] self.cv3: has the same structure as self.cv2, and also contains a RepNCSP module and a 3x1 convolution module.

[0063] self.cv4: A 1x1 convolution that converts the merged features from C3+2*C4 channels to C2 channels.

[0064] The RepNCSPELAN4 module operates by forward propagating the input feature x, first performing a 1x1 convolution with self.cv1. The result is then processed in two branches, one of which is processed by self.cv2 and self.cv3. Multiple convolutions are performed within each RepNCSP module. Finally, the features processed by self.cv2 and self.cv3 are combined with the original C3 channel features from the other branch, and a 1x1 convolution is performed with self.cv4 to output the final features. This design helps to better aggregate multi-scale features and enhance the model's ability to recognize objects of different sizes.

[0065] S33, add auxiliary detection head DetectAux to the detection head part of the Head end: the main detection head is mainly used for the detection of large targets, and the detection effect of small targets is not ideal, such as Figure 2 As shown, this embodiment uses three DetectAux modules on the Head end as auxiliary detection heads to assist the main detection head in completing the detection task of small targets. The auxiliary detection head is usually located in the middle layer of the network, or the auxiliary detection head is added to the head part of the model and together with the main detection head constitutes the output layer of the YOLOv8 model. The DetectAux module plays an additional supervisory role in the model training process, which helps to improve the stability and generalization ability of the network, especially in the early stage of model training or when dealing with some complex scenes. It can capture feature information of different scales or levels, which helps to more comprehensively understand the targets in the image in the YOLOv8 code implementation.

[0066] The loss function of the auxiliary detection head DetectAux is similar to the loss function of the main detection head, which mainly consists of two parts: classification loss and bounding box regression loss. Specifically, the loss function can be expressed as:

[0067] Loss=BCEcls(d1)+BCEcls(d2)+CIoU(d1)+CIoU(d2)+DFL(d1)+DFL(d2)

[0068] Among them, BCEcls represents the binary cross entropy loss, which is used for classification tasks; CIoU represents the CIoU loss, which is used for bounding box regression tasks; DFL represents the distribution focus loss, which is used for bounding box regression tasks; d1 and d2 represent the outputs of the main detection head and auxiliary detection head, respectively.

[0069] The classification loss (BCEcls) uses a binary cross entropy loss function to measure the difference between the classification model prediction result and the true label. The formula is expressed as:

[0070]

[0071] Among them, y i is the true label of the i-th sample, is the predicted label of the i-th sample, and N is the number of samples.

[0072] Bounding box regression loss (CIoU) is an improved IoU loss that takes into account the shape and size of the bounding box. The formula is expressed as:

[0073]

[0074] Among them, IoU(b,b * ) is the predicted box b and the real box b * The intersection-and-union ratio of p 2 (b,b * ) is the square of the Euclidean distance between the center point of the predicted box and the real box; c is the diagonal length of the minimum closure containing the predicted box and the real box; ar(b) and ar(b * ) are the aspect ratios of the predicted box and the true box respectively; α is an adjustment parameter used to balance the weights of different parts.

[0075] The distribution focus loss is used for bounding box regression tasks. It uses the distribution focus mechanism to reduce the weight of easy-to-classify samples and increase the weight of difficult-to-classify samples. The formula of the distribution focus loss is:

[0076]

[0077] Where y is the true label, y^ is the predicted label, and N is the number of samples.

[0078] This results in the YOLOv8-ARD model, which also includes a backbone for feature extraction, a neck for feature fusion, and a head for prediction. The backbone of the YOLOv8-ARD model includes the CBS module, the Neck module, the RepNCSPELAN4 module, and the SPPF module. The CBS module consists of a convolutional layer, a normalization layer, and a SiLU activation function layer, while the RepNCSPELAN4 module consists of a CBS module, a RepNCSP module, a split layer, and several Bottleneck layers. RepNCSPELAN4's backbone module, which replaces the RepNCSPELAN4 module, enhances learning of image features for areas such as defects, spontaneous explosions, and flashovers in power transmission lines, improving target detection accuracy and enhancing the model's generalization. The Neck side of the YOLOv8-ARD model includes the CBS module, RepNCSPELAN4 module, ADown module, Upsample layer, and BiFPN_Concat module. The Upsample layer performs upsampling operations to facilitate feature fusion. The Neck side, combined with the BiFPN structure, performs multi-scale feature fusion, which improves the YOLOv8-ARD model's accuracy for small object detection while reducing network complexity and redundant computation. The Head side of the YOLOv8-ARD model uses the auxiliary detection head DetectAux, which can capture feature information at different scales or levels, facilitating a more comprehensive understanding of objects in the image.

[0079] S4, training YOLOv8-ARD model: input the training set in the labeled dataset into the YOLOv8-ARD model for training, obtain the YOLOv8-ARD.pt weight file, and then obtain the YOLOv8-ARD model. After the training is completed, use the validation set for verification.

[0080] S5, using the YOLOv8-ARD model to identify foreign objects invading the transmission line: Use the trained YOLOv8-ARD model to detect and locate foreign objects in the images of foreign objects invading the transmission line to be identified, and promptly repair the transmission line where the faulty transmission line is located.

[0081] The framework used in this embodiment is Pytorch 1.12, the Python version is 3.8, the CUDA version is 12.1, the training system is Windows 11, and the graphics card model used for training is NVIDIA GeForce RTX 4060 16G.

[0082] The YOLOv8 model and the YOLOv8-ARD model are tested on the same dataset. Figure 6This is a comparison chart of the detection results of the YOLOv8 model and the YOLOv8-ARD model for insulator equipment detection in transmission lines where there are artificial high-altitude foreign objects. Figure 6 It can be seen that the probability of detecting the kite is 0.28 and 0.83 respectively. Figure 7 The following figure compares the detection results of the YOLOv8-ARD model with those of the YOLOv5, YOLOv6, and YOLOv8 models. The results show that the YOLOv8-ARD network model of this embodiment improves accuracy by 2.2%, regression rate by 1.7%, map50 by 2.2%, and map50-95 by 1.3% compared to the original YOLOv8 model. This shows that compared to the YOLOv8 model, the YOLOv8-ARD model of this embodiment has stronger feature extraction capabilities, performs better in distinguishing similar features, more thoroughly integrates multi-scale features, and is more capable of detecting small targets, with lower missed detection and false detection rates. It achieves excellent detection results even in complex outdoor environments, exhibits good robustness and generalization, and can meet the requirements of unmanned inspections. This demonstrates that the improved YOLOv8-ARD network model of this embodiment improves both accuracy and regression rate in detecting insulator equipment and its defects. The use of the YOLOv8-ARD network model for transmission line foreign object intrusion detection improves the accuracy of transmission line foreign object intrusion detection, which is crucial to ensuring the safe and stable operation of the power grid.

[0083] The embodiments described above are merely descriptions of preferred implementations of the present invention and are not intended to limit the scope of the present invention. Without departing from the concept of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should fall within the scope of protection determined by the claims of the present invention.

Claims

1. A method for detecting foreign objects in power transmission lines based on multi-scale fusion and enhanced adaptation, characterized in that: The following detection steps are included: S1, obtain the image dataset of foreign objects intruding on the transmission line and screen and classify it; S2, manually annotate the screened foreign body invasion images and divide them into training set and validation set; S3: Build the YOLOv8-ARD model based on the original YOLOv8 model. The original YOLOv8 model includes the Backbone end for feature extraction, the Neck end for feature fusion, and the Head end for prediction. The process of building the YOLOv8-ARD model includes: The downsampling layers on the Backbone and Head sides of the YOLOv8 model are replaced with ADown modules. The C2F module in the YOLOv8 model is replaced with the RepNCSPELAN4 module. An auxiliary detection head is added to the detection head on the Head side to assist the main detection head in completing the detection task of small objects. S4, training YOLOv8-ARD model: input the training set in the labeled dataset into the YOLOv8-ARD model for training to obtain the YOLOv8-ARD.pt weight file; S5 uses the trained YOLOv8-ARD model to detect and locate foreign objects in the transmission line foreign object intrusion image to be identified.

2. The method for detecting foreign objects in power transmission lines based on multi-scale fusion and enhanced adaptation according to claim 1, characterized in that: The implementation process of step S1 is as follows: S11, using unmanned inspection equipment to collect images of foreign objects intruding on the transmission line; S12, manually screening the acquired images and obtaining an image data set, wherein during the manual screening, intact images of foreign objects intruding the transmission line are removed, and images of foreign objects intruding the transmission line with obstructions are retained; S13, classifying the screened transmission line foreign body images.

3. The method for detecting foreign objects in power transmission lines based on multi-scale fusion and enhanced adaptation according to claim 1, characterized in that: In S2, the process of manually labeling the screened foreign body intrusion images and dividing them into a training set and a validation set is as follows: S21, using an image annotation tool and selecting the YOLO annotation format to annotate the classified foreign body intrusion image; S22, save each label name and its corresponding bounding box into a .txt file, the saved file forms a dataset, and use a Python script to randomly divide the dataset into a training set and a validation set in a ratio of 8:

2.

4. The method for detecting foreign objects in power transmission lines based on multi-scale fusion and enhanced adaptation according to claim 1, characterized in that: In S3, the ADown module includes two branches, namely cv1 and cv2. cv1 uses a 3x3 convolution kernel for downsampling, and cv2 first performs 3x3 maximum pooling and then uses a 1x1 convolution kernel for downsampling. The ADown module processes the data set as follows: in the forward propagation, the feature map is input into the ADown module, the feature map is downsampled by average pooling, and the downsampled feature map is divided into two parts, one part is processed by cv1 and the other part is processed by cv2, and finally the two parts are spliced ​​together.

5. The method for detecting foreign objects in power transmission lines based on multi-scale fusion and enhanced adaptation according to claim 1, characterized in that: In S3, the parameters of the RepNCSPELAN4 module are C1, C2, C3, C4 and C5, where C1 is the number of input channels, C2 is the number of output channels, C3 is the number of intermediate feature channels, usually half of C2, C4 is the number of channels in the RepNCSP module, and C5 is the number of times the RepNCSP module is repeated, which defaults to 1; the structure of the RepNCSPELAN4 module includes self.cv1, self.cv2, self.cv3 and self.cv4, self.cv1 input first After a 1x1 convolution, the input features are converted from C1 channels to C3 channels; then the feature map is divided into two parts; self.cv2 contains a RepNCSP module and a 3x1 convolution module. The RepNCSP module converts the features of the C3 / 2 channels to the C4 channels, and then keeps the number of channels unchanged through 3x1 convolution; self.cv3 contains a RepNCSP module and a 3x1 convolution module; self.cv4 includes a 1x1 convolution, which converts the merged features from C3+2*C4 channels to C2 channels.

6. The method for detecting foreign objects in power transmission lines based on multi-scale fusion and enhanced adaptation according to claim 1, characterized in that: In the S3, the Head end uses three DetectAux modules as auxiliary detection heads.

7. The method for detecting foreign objects in power transmission lines based on multi-scale fusion and enhanced adaptation according to claim 6, characterized in that: In the process of training the YOLOv8-ARD model, the loss function of the auxiliary detection head DetectAux includes classification loss and bounding box regression loss; The loss function is expressed as: Loss=BCEcls(d1)+BCEcls(d2)+CIoU(d1)+CIoU(d2)+DFL(d1)+DFL(d2) Among them, BCEcls represents binary cross entropy loss, which is used for classification tasks; CIoU represents CIoU loss, which is used for bounding box regression tasks; DFL represents distribution focus loss, which is used for bounding box regression tasks; d1 and d2 represent the outputs of the main detection head and auxiliary detection head, respectively; Classification loss is a function used to measure the difference between the classification model prediction results and the true label. The formula is expressed as: Among them, y i is the true label of the i-th sample, is the predicted label of the i-th sample, and N is the number of samples. The formula for bounding box regression loss is expressed as: Among them, IoU(b,b * ) is the predicted box b and the real box b * The intersection-and-union ratio of p 2 (b,b * ) is the square of the Euclidean distance between the center point of the predicted box and the real box; c is the diagonal length of the minimum closure containing the predicted box and the real box; ar(b) and ar(b * ) are the aspect ratios of the predicted box and the true box respectively; α is an adjustment parameter used to balance the weights of different parts; Distribution focus loss is used for bounding box regression tasks. The formula of distribution focus loss is: Where y is the true label, y^ is the predicted label, and N is the number of samples.

Citation Information

Cited By

  • Lightweight power transmission line insulator defect detection method and system

    CN122223596A