Power transmission line foreign matter detection method based on deep learning

By improving the Yolov8 algorithm, a TL_Yolo model was established, and a full-dimensional dynamic convolution and an improved BiFPN network were used, combined with the head network of attention mechanism, and the problems of missed detection, false detection and inaccurate detection of small objects in the transmission line were solved, achieving more efficient and accurate detection effects.

CN120198341AInactive Publication Date: 2025-06-24NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311775292.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems such as missed detection, missed detection and inaccurate detection of small targets in the detection of foreign objects in transmission lines.

Method used

Using a deep learning-based method, the TL_Yolo model is established by improving the Yolov8 object detection algorithm, and the target features are extracted and fused by using full-dimensional dynamic convolution and improved BiFPN network, and the head network with attention mechanism is introduced for prediction output.

Benefits of technology

It improves the accuracy and efficiency of foreign object detection in transmission lines, reduces missed and missed detection, and can detect small targets more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198341A_ABST
    Figure CN120198341A_ABST
Patent Text Reader

Abstract

The invention discloses a transmission line foreign matter detection algorithm based on deep learning. The transmission line foreign matter detection algorithm comprises the following steps: selecting a mainstream Yolov8 target detection algorithm as a basic framework; preprocessing the target image; target features are efficiently extracted through a Backbone network based on full-dimensional dynamic convolution; carrying out deep fusion on the extracted target features through a Neck network based on the improved BiFPN; the foreign matter of the power transmission line is accurately predicted by introducing the Head network of the attention mechanism; and designing a loss function to calculate classification and regression loss. According to the invention, through the methods of image preprocessing, feature extraction network improvement, feature fusion capability enhancement and the like, a prediction output feature map with higher precision is obtained, and the problems of missing detection, false detection and inaccurate small target detection of foreign matter detection of the power transmission line in the prior art are solved. And a more efficient and accurate power transmission line foreign matter detection method is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a foreign object detection method for transmission lines based on deep learning. Background Art

[0002] As a transmission carrier of electric energy, the transmission line is an important part of the power system. The transmission line is often affected by climate change and foreign objects (such as bird nests, kites, plastic bags, balloons, etc.). Therefore, in order to ensure the safe and stable operation of the power system, it is very necessary to detect foreign objects on the transmission line. In recent years, the construction speed of China's power grid has gradually accelerated, and the areas where transmission lines are erected have also shown a trend of densification and complexity, which has also increased the difficulty of maintenance work for transmission lines by staff.

[0003] Currently, the methods for detecting foreign objects on transmission lines are mainly divided into two categories: one is manual inspection, relying on inspectors to walk along the transmission line on foot or with the help of a helicopter, directly observing whether there are foreign objects with the naked eye, and recording the inspection results in a paper medium; the other is intelligent inspection, collecting images or videos through relevant equipment such as drones, and using artificial intelligence algorithms to analyze and judge foreign objects.

[0004] The method of manual inspection has problems such as high labor intensity, low work efficiency, poor reliability, and there is also a certain degree of danger; in the method of intelligent inspection, images are obtained by drones. However, when facing a large amount of data, staff often have situations of false detection or missed detection, and the manual detection time is long, and foreign objects on the transmission line cannot be found in a timely and accurate manner, which is not conducive to the inspection of transmission lines. Summary of the Invention

[0005] The purpose of the present invention is to provide a foreign object detection method for transmission lines based on deep learning to solve the technical problems of missed detection, false detection, and inaccurate detection of small targets in the current foreign object detection of transmission lines.

[0006] To solve the above technical problems, the specific technical solution of the present invention is as follows:

[0007] A foreign object detection method for transmission lines based on deep learning, comprising the following steps:

[0008] Step 1, select the Yolov8 object detection algorithm as the framework;

[0009] Step 2, preprocess the target image;

[0010] Preprocess the data at the data input end, uniformly adjust the size of the input picture to the input size of the network, and adopt the Mosaic data augmentation technology to balance the number of small, medium, and large targets in the data set;

[0011] Step 3. Improve the Yolov8 object detection algorithm to establish the TL_Yolo model, which specifically includes the following steps:

[0012] Step 3.1. Extract target features through a backbone network based on Omni-Dimensional Dynamic Convolution.

[0013] The Yolov8 object detection algorithm uses the C2f (CSPDarknet53 to 2-Stage FPN) structure to adjust different numbers of channels for images with different scales of input.

[0014] In the Backbone network of TL_Yolo, use Omni-Dimensional Dynamic Convolution to replace the standard convolution in the bottleneck structure to obtain a bottleneck structure based on Omni-Dimensional Dynamic Convolution (Bottleneck_ODConv); integrate Bottleneck_ODConv into the C2f module to obtain a C2f module based on Omni-Dimensional Dynamic Convolution (C2f_ODConv). Omni-Dimensional Dynamic Convolution introduces a multi-dimensional attention mechanism, which has a parallel strategy for learning different information of the convolutional kernel in all four dimensions in the kernel space. The four dimensions include the convolutional kernel spatial size, input channels, output channels, and the number of convolutional kernels. The specific definition of the Omni-Dimensional Dynamic Convolution operation is as follows:

[0015] y = (α w1 ⊙ α f1 ⊙ α c1 ⊙ α s1 ⊙ W1 +... + α wn ⊙ α fn ⊙ α cn ⊙ α sn ⊙ W n ) * x (1)

[0016] In formula (1), α wi , α si , α ci , α fi represent the attention calculated by the convolutional kernel W i along the dimension dimension, input channel dimension, output channel dimension, and spatial dimension respectively. x represents the input information, and y represents the output information.

[0017] Step 3.2: Fuse the extracted target features through the Neck network based on the improved BiFPN;

[0018] The Neck network of TL_Yolo is located in the middle position between the Backbone network and the Head output end, and fully fuses the features extracted by the Backbone network to enhance the diversity and robustness of the features; TL_Yolo adopts the Feature Pyramid Networks (FPN) and Path Aggregation Network (PAN) structures. The FPN structure samples from top to bottom, enabling the bottom feature map to contain strong semantic information of the image; the PAN structure samples from bottom to top, enabling the top feature to contain the location information of the image; the combination of FPN and PAN aggregates the parameters from different backbone layers, and introduces the weighted bidirectional feature pyramid BiFPN to strengthen the bottom information of the feature map, enabling information fusion of feature maps of different scales; in BiFPN, the weighted fusion method is used, and the one-shot aggregation network (VoVGSCSP) structure is selected to replace the original C2f structure as the node;

[0019] Step 3.3: Predict and output the feature map through the Head network with the attention mechanism introduced;

[0020] The Head network of TL_Yolo is used to output the object detection results; it adopts the Anchor-Free idea and also uses the Decoupled-Head, and outputs the classification and regression results through two heads respectively; the Dynamic head based on the attention mechanism is introduced, and the three feature maps obtained by the feature fusion of the Neck network are input into the Dynamic head for processing, and then the processed feature maps are used to predict the results;

[0021] Step 4: Design a loss function to calculate the classification and regression losses;

[0022] The loss function calculation includes two parts: the classification branch and the regression branch. The classification branch uses the improved cross-entropy loss (VFL Loss), and the regression branch uses the Distribution Focal Loss and the CIoU loss L CIOU 。

[0023] Furthermore, Step 2 specifically includes the following steps:

[0024] Step 2.1: Randomly select four pictures from the data;

[0025] Step 2.2: Perform left - right flipping (flip the original image left - right), size scaling (scale the size of the original image), and color gamut change (change the brightness, saturation, and hue of the original image) operations on the four pictures respectively;

[0026] Step 2.3: Piece together and combine the pictures in sequence.

[0027] Furthermore, the extraction of the target feature in Step 3.1 includes the following steps:

[0028] Step 3.1.1: Continuously use two 3*3 convolutions to downsample the feature map P1 by 4 times to obtain the feature map P2;

[0029] Step 3.1.2: Use the C2f_ODConv module to collect the gradient flow information after being downsampled by 4 times for the next sampling work;

[0030] Step 3.1.3: Successively downsample the feature map by 8 times, 16 times, and 32 times according to Steps 3.1 - 3.2 to obtain the corresponding feature maps P3, P4, and P5;

[0031] Step 3.1.4: Concatenate the features of the same feature map at different scales together through the SPPF module.

[0032] Furthermore, Step 3.2 specifically includes the following steps:

[0033] Step 3.2.1: Perform convolution operations on the feature maps P3, P4, and P5 extracted by the Backbone network to unify the number of channels of the feature maps;

[0034] Step 3.2.2: Upsample the feature map P5, fuse it with P4 using the weight fusion method, and obtain the feature map P54 at this time through the node VoVGSCSP;

[0035] Step 3.2.3: Upsample the feature map P54, fuse it with P3 using the weight fusion method, and obtain the feature map P*3 at this time through the node VoVGSCSP. Thus, the weighted bidirectional feature pyramid BiFPN completes the top - down feature fusion work;

[0036] Step 3.2.4: Downsample the feature map P2, fuse it with the feature maps P3 and P54 through the node VoVGSCSP to obtain the feature map P23 at this time;

[0037] Step 3.2.5: Downsample the feature map P23, fuse it with the feature maps P4 and P54 through the node VoVGSCSP to obtain the feature map P34 at this time;

[0038] Step 3.2.6: Downsample the feature map P34 and fuse it with the feature map P5 to obtain the feature map P*5 through the node VoVGSCSP. Thus, BiFPN completes the bottom-up feature fusion work;

[0039] Step 3.2.7: After Steps 3.2.1 - 3.2.6, a feature map that fuses high-level semantic information and low-level position information is obtained.

[0040] Further, Step 3.3 specifically includes the following steps:

[0041] Before the output of the Head network, introduce the object detection head Dynamic head based on the attention mechanism. Apply the attention mechanism from three different perspectives: scale perception, spatial perception, and task perception, to improve the expression ability of the object detection head of the TL_Yolo model without increasing the computational cost; Dynamic head is the superposition of three kinds of attention, and the stacking of the three kinds of attention is called a DyHead block, where π L , π S , π C represent scale perception attention, spatial perception attention, and task perception attention respectively. The specific composition of Dynamic head is:

[0042] The scale perception module consists of average pooling + 1x1 convolution + ReLU + sigmoid function;

[0043] The spatial perception module consists of index + 3x3 convolution + sigmoid + offset; where the index operation uses different deformable convolutions for feature maps of different sizes; offset is a standard 3x3 convolution;

[0044] The task perception module consists of average pooling + fully connected layer + ReLU + fully connected layer + Normalize; where Normalize is implemented through sigmoid, and the specific formula is:

[0045]

[0046] The combination of max and min operations ensures that the output value of sigmoid is in the range of 0 to 1, and controls the ON / OFF of different feature map channels according to different tasks, so as to achieve task perception; Input the feature maps P3, P4, and P5 after feature fusion into the DyHead block to perform scale perception, spatial perception, and task perception operations in sequence, and then input the feature maps into the Decoupled-Head for prediction.

[0047] Furthermore, the Loss calculation of the TL_Yolo model includes two parts. First, the VFL Loss is calculated for the classification branch, and its calculation formula is:

[0048]

[0049] where p is the predicted value, q is the true label value, and η and γ are weight parameters; for the regression branch, the Distribution Focal Loss and CIoU Loss are calculated. Among them, the Distribution Focal Loss uses the cross-entropy function to optimize the probabilities of the two positions on the left and right near the label, enabling the model to quickly focus on the values near the label. Its calculation formula is

[0050] DFL(S i ,S i+1 )=-((y i+1 -y)log(S i )+(y-y i )log(S i+1 )) (4)

[0051] where S i is the sigmod output of the network, i and i + 1 are the interval orders, and y is the label value; the CIoU loss L CIOU takes into account the distance, overlap rate, scale, and penalty term between the target and the anchor, making the target box regression more stable. Its calculation formula is:

[0052]

[0053] where α is the weight function, ν is used to measure the consistency of the aspect ratio, b represents the predicted box, and b gt represents the true box, ρ 2 (b,b gt ) represents the Euclidean distance between the centers of the predicted box and the true box, c represents the diagonal distance of the smallest closed region that can contain both the predicted box and the true box, and IOU is the intersection over union of the predicted box and the true box, that is

[0054]

[0055] Transmission Line Yolo (TL_Yolo) has made several improvements compared to Yolov8:

[0056] 1) Efficiently extract target features through a Backbone network based on Omni-Dimensional Dynamic Convolution; 2) Deeply fuse the extracted target features through a Neck network based on improved BiFPN; 3) Accurately predict foreign objects on transmission lines through a Head network introducing an attention mechanism.

[0057] A method for detecting foreign objects on transmission lines based on deep learning according to the present invention has the following advantages:

[0058] Based on the object detection technology of deep learning, compared with common object detection models, through methods such as image preprocessing, improving the feature extraction network, and enhancing the feature fusion ability, a prediction output feature map with higher accuracy is obtained, overcoming the problems of missed detection, false detection, and inaccurate detection of small targets in the prior art for detecting foreign objects on transmission lines, and obtaining a more efficient and accurate method for detecting foreign objects on transmission lines.

[0059] By optimizing the traditional Yolov8 algorithm, the present invention solves the problems of missed detection and inaccurate detection of small targets in the original method. In addition, with the improvement of the computing performance of computer hardware, especially graphics processing units (GPUs), the detection speed and accuracy of images have been significantly improved. And the real-time detection of targets completes the integration of "inspection and detection", further improving the inspection efficiency of abnormal conditions of transmission lines and further promoting the development of power grid intelligence. Brief Description of the Drawings

[0060] Figure 1 It is the overall flowchart of the present invention;

[0061] Figure 2 It is the model structure diagram of the present invention;

[0062] Figure 3 It is the flowchart of the Backbone network of the present invention for extracting image features;

[0063] Figure 4 It is the structure diagram of C2f_ODConv of the present invention;

[0064] Figure 5 It is the structure diagram of Bottleneck_ODConv of the present invention;

[0065] Figure 6 It is the working principle diagram of Omni-Dimensional Dynamic Convolution of the present invention;

[0066] Figure 7 It is the flowchart of the Neck network of the present invention for feature fusion;

[0067] Figure 8 This is the diagram of the weight Fusion method of the present invention;

[0068] Figure 9 This is the structure diagram of GSBottleneck and VoVGSCSP of the present invention;

[0069] Figure 10 This is the flowchart of the output prediction of the Head network of the present invention;

[0070] Figure 11 This is the structure diagram of the Dynamic block of the present invention;

[0071] Figure 12 This is the comparison diagram of the detection results between the right side of the present invention and the left side of the traditional Yolov8. Detailed implementation manners

[0072] To better understand the purpose, structure and function of the present invention, the following further describes in detail a foreign object detection method for transmission lines based on deep learning according to the present invention with reference to the accompanying drawings. Figure 1 This is the overall flowchart of the present invention.

[0073] Step 1: Select the Yolov8 object detection algorithm as the framework;

[0074] Step 2: Preprocess the target image;

[0075] Preprocess the data at the data input end, uniformly adjust the size of the input pictures to 640*640, and adopt the Mosaic data augmentation technology to improve the problem of uneven numbers of small, medium and large targets in the dataset. The main steps are as follows:

[0076] 1) Randomly select four pictures from the data;

[0077] 2) Perform operations such as flipping (flipping the original picture left and right), scaling (scaling the size of the original picture), and color gamut change (changing the brightness, saturation and hue of the original picture) on the four pictures respectively;

[0078] 3) Stitch and combine the pictures in order.

[0079] Step 3: Extract the target features through the Backbone of the Omni-Dimensional Dynamic Convolution-based backbone network;

[0080] The Yolov8 object detection algorithm adopts the C2f structure and adjusts different numbers of channels for pictures with different scale inputs.

[0081] In the Backbone network of TL_Yolo, the standard convolution in the bottleneck is replaced by the Omni-Dimensional Dynamic Convolution to obtain a bottleneck structure Bottleneck_ODConv based on the Omni-Dimensional Dynamic Convolution. The specific calculation process is as Figure 2 , obtaining a new structure Bottleneck_ODConv, as Figure 3 shown; integrating Bottleneck_ODConv into the C2f module to obtain a new module C2f_ODConv, as Figure 4 shown. The Omni-Dimensional Dynamic Convolution introduces a multi-dimensional attention mechanism, which has a parallel strategy for learning different information of the convolution kernel in all four dimensions (convolution kernel spatial size, input channels, output channels, number of convolution kernels) in the kernel space. The specific definition is as follows:

[0082] y = (α w1 ⊙α f1 ⊙α c1 ⊙α s1 ⊙W1 +... + α wn ⊙α fn ⊙α cn ⊙α sn ⊙W n ) * x (1)

[0083] In formula (1), α wi ∈R represents the attention α i along the number of convolution kernels W obtained through the attention function, while α wi , α si ∈R k*k , respectively represent the attention calculated along the spatial dimension, input channel dimension, and output channel dimension of the convolution kernel W i . x represents the input information, and y represents the output information. Figure 5 This is the flowchart for the Backbone network of the present invention to extract image features. The main process is as follows:

[0084] 1) Continuously use two 3*3 convolutions to downsample the feature map by 4 times to obtain the feature map P2;

[0085] 2) Use the C2f_ODConv module to collect the gradient flow information after downsampling by 4 times for the next sampling work;

[0086] 3) According to the above steps, the feature maps are downsampled by 8 times, 16 times, and 32 times to obtain the corresponding feature maps P3, P4, and P5;

[0087] 4) The features of the same feature map at different scales are spliced ​​together through the SPPF module.

[0088] Step 4: The extracted target features are fused through the Neck network based on the improved BiFPN;

[0089] The Neck network of TL_Yolo is located in the middle of the Backbone network and the Head output end, and fully integrates the features extracted by the Backbone network; TL_Yolo adopts the feature pyramid network FPN and the path aggregation network PAN structure. The FPN structure samples from top to bottom, so that the bottom feature map contains strong semantic information of the image; the PAN structure samples from bottom to top, so that the top feature contains the image position information; FPN is combined with PAN to aggregate the parameters from different backbone layers, and the weighted bidirectional feature pyramid BiFPN is introduced to strengthen the bottom information of the feature map, so that the feature maps of different scales can be integrated; the weighted fusion method is used in BiFPN, and the one-time aggregation network VoVGSCSP structure is selected to replace the original C2f structure as the node, such as Figure 6 , 7 As shown. The weighted bidirectional feature pyramid BiFPN structure is as follows Figure 8 As shown in the figure, the specific fusion process is:

[0090] 1) Perform convolution operations on the feature maps P3, P4, and P5 extracted by the Backbone network to unify the number of channels of the feature maps;

[0091] 2) Upsample the feature map P5 and fuse it with P4 using the weight fusion method, and obtain the feature map P54 at this time through the node VoVGSCSP;

[0092] 3) Upsample the feature map P54 and fuse it with P3 using the weight fusion method. Then, the feature map P*3 is obtained through the node VoVGSCSP. At this point, BiFPN completes the top-down feature fusion work.

[0093] 4) Downsample the feature map P2, fuse it with the feature map P3 and the feature map P54 through the node VoVGSCSP to obtain the feature map P23 at this time;

[0094] 5) Downsample the feature map P23 and fuse it with the feature maps P4 and P54 through the node VoVGSCSP to obtain the feature map P34 at this time;

[0095] 6) Downsample the feature map P34 and fuse it with the feature map P5 to obtain the feature map P*5 through the node VoVGSCSP. So far, BiFPN has completed the bottom-up feature fusion work;

[0096] 7) After the above process, a feature map that integrates high-level semantic information and low-level position information is obtained.

[0097] Step 5: Predict and output the feature map by introducing the Head network with attention mechanism;

[0098] The Head network of TL_Yolo is used to output the target detection results. It adopts the idea of ​​anchor-free frame and uses the decoupled-Head to output the classification and regression results respectively through two heads. It introduces the target detection head Dynamic head based on the attention mechanism, and inputs the three feature maps obtained by the Neck network feature fusion into the Dynamic head for processing, and then uses the processed feature maps to predict the results.

[0099] Before the output of the Head network, the target detection head Dynamic head based on the attention mechanism is introduced. The attention mechanism is used from three different angles: scale perception, space perception, and task perception. The expression ability of the model target detection head is significantly improved without increasing the amount of calculation. Dynamic head is the superposition of three types of attention. The stacking of three types of attention is called a DyHead block, where π L ,π S ,π C They represent scale-aware attention, space-aware attention, and task-aware attention, respectively. Figure 9 Its specific composition is:

[0100] 1) The scale-aware module consists of average pooling + 1x1 convolution + ReLU + sigmoid function;

[0101] 2) The spatial perception module consists of index+3x3 convolution+sigmoid+offset; the index operation uses different deformable convolutions for feature maps of different sizes; the offset is a standard 3x3 convolution.

[0102] 3) The task perception module consists of average pooling + fully connected layer + ReLU + fully connected layer + Normalize; Normalize is implemented by sigmoid, and the specific formula is:

[0103]

[0104] The combination of max and min operations ensures that the output value of the sigmoid is within the range of 0 to 1; the on / off of different feature map channels is controlled according to different tasks to achieve task awareness. The feature maps P3, P4, and P5 after feature fusion are input into the DyHead block to perform scale awareness, spatial awareness, and task awareness operations in sequence, and then the feature maps are input into the Decoupled-Head for prediction, as Figure 10 shown.

[0105] Step 6: Design a loss function to calculate classification and regression losses;

[0106] The loss function calculation includes two parts: the classification branch uses the improved cross-entropy loss VFL Loss, and the regression branch uses the Distribution Focal Loss and the CIoU loss L CIOU .

[0107] The Loss calculation of the model includes two parts. First, the VFL Loss is calculated for the regression branch, and its calculation formula is:

[0108]

[0109] where p is the predicted value, q is the true label value, and α and γ are weight parameters. For the regression branch, the Distribution Focal Loss and the CIoU Loss are calculated. The design idea of the Distribution Focal Loss is to use the cross-entropy function to optimize the probabilities of the two positions on the left and right near the label, so that the network can quickly focus on the values near the label. Its calculation formula is

[0110] DFL(S i ,S i+1 ) = -((y i+1 -y)log(S i )+(y-y i )log(S i+1 )) (4)

[0111] where S i is the sigmod output of the network, y i and y i+1 are the interval orders, and y is the label value. The CIoU Loss takes into account the distance, overlap rate, scale, and penalty term between the target and the anchor, making the target box regression more stable. Its calculation formula is:

[0112]

[0113] where α is a weight function, ν is used to measure the aspect ratio consistency, b represents the predicted bounding box, and b gt represents the ground truth bounding box, ρ 2 (b, b gt ) represents the Euclidean distance between the centers of the predicted bounding box and the ground truth bounding box, c represents the diagonal distance of the smallest closed region that can simultaneously contain the predicted bounding box and the ground truth bounding box, and IOU is the intersection over union of the predicted bounding box and the ground truth bounding box, that is

[0114]

[0115] The loss value of the model can be obtained by weighting the three Losses with a certain weight ratio.

[0116] Figure 11 This is the Dynamic block structure diagram of the present invention. Figure 12 This is a comparison chart of the detection results of the present invention (right side) and traditional Yolov8 (left side). The method of the present invention solves the problems of missed detection of bird nests, misdetection of kites, and low detection accuracy of small balloon targets existing in traditional Yolov8.

[0117] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that, without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. A foreign object detection method for transmission lines based on deep learning, characterized in that, It includes the following steps: Step 1: Select the Yolov8 object detection algorithm as the framework; Step 2: Preprocess the target image; Preprocess the data at the data input end, uniformly adjust the size of the input image to the input size of the network, and use the Mosaic data augmentation technique to balance the number of small, medium, and large targets in the dataset; Step 3: Improve the Yolov8 object detection algorithm to establish the TL_Yolo model, which specifically includes the following steps: Step 3.1: Extract the target features through the Backbone network based on the Omni-Dimensional Dynamic Convolution; The Yolov8 object detection algorithm uses the C2f structure and adjusts different numbers of channels for images with different scales of input; In the Backbone network of TL_Yolo, use the Omni-Dimensional Dynamic Convolution to replace the standard convolution in the bottleneck structure to obtain a Bottleneck_ODConv based on the Omni-Dimensional Dynamic Convolution; integrate Bottleneck_ODConv into the C2f module to obtain a C2f module based on the Omni-Dimensional Dynamic Convolution; the Omni-Dimensional Dynamic Convolution introduces a multi-dimensional attention mechanism, which has a parallel strategy for learning different information of the convolution kernel in all four dimensions in the kernel space. The four dimensions include the spatial size of the convolution kernel, the input channels, the output channels, and the number of convolution kernels. The specific definition of the Omni-Dimensional Dynamic Convolution operation is as follows: y = (α w1 ⊙ α f1 ⊙ α c1 ⊙ α s1 ⊙ W1 +... + α wn ⊙ α fn ⊙ α cn ⊙ α sn ⊙ W n ) * x (1) α in formula (1) wi , α si , α ci , α fi represent the attention calculated by the convolutional kernel W i along the dimension dimension, input channel dimension, output channel dimension, and spatial dimension respectively. x represents the input information, and y represents the output information; Step 3.2: Fuse the extracted target features through the Neck network based on the improved BiFPN; The Neck network of TL_Yolo is located in the middle position between the Backbone network and the Head output end to fully fuse the features extracted by the Backbone network; TL_Yolo adopts the Feature Pyramid Network (FPN) and the Path Aggregation Network (PAN) structure. The FPN structure samples from top to bottom, so that the bottom feature map contains strong semantic information of the image; the PAN structure samples from bottom to top, so that the top feature contains the position information of the image; the FPN and PAN are combined to aggregate the parameters from different backbone layers, and the weighted bidirectional feature pyramid (BiFPN) is introduced to strengthen the bottom information of the feature map, enabling information fusion of feature maps of different scales; in the BiFPN, the weighted fusion method is used, and the one-time aggregation network VoVGSCSP structure is selected to replace the original C2f structure as the node; Step 3.3: Predict and output the feature map through the Head network introducing the attention mechanism; The Head network of TL_Yolo is used to output the object detection results. The idea of Anchor-Free is adopted, and the Decoupled-Head is used at the same time. The classification and regression results are output through two heads respectively. The Dynamic head of object detection based on the attention mechanism is introduced. The three feature maps obtained by the feature fusion of the Neck network are input into the Dynamic head for processing, and then the processed feature maps are used to predict the results. Step 4: Design a loss function to calculate the classification and regression losses. The loss function calculation includes two parts: the classification branch and the regression branch. The improved cross-entropy loss VFLLoss is adopted in the classification branch, and the distribution focal loss and the CIoU loss L are used in the regression branch CIOU .

2. The foreign object detection method for transmission lines based on deep learning according to claim 1, characterized in that Step 2 specifically includes the following steps: Step 2.1: Randomly select four pictures from the data. Step 2.2: Perform operations of left-right flipping, size scaling, and color gamut change on the four pictures respectively. Step 2.3: Stitch and combine the pictures in order.

3. The foreign object detection method for transmission lines based on deep learning according to claim 2, characterized in that, The extraction of the target features in Step 3.1 includes the following steps: Step 3.1.1: Continuously use two 3*3 convolutions to downsample the feature map P1 by 4 times to obtain the feature map P2. Step 3.1.2: Use the C2f_ODConv module to collect the gradient flow information after downsampling by 4 times for the next downsampling work. Step 3.1.3: Downsample the feature map by 8 times, 16 times, and 32 times in sequence according to Step 3.1 - Step 3.2 to obtain the corresponding feature maps P3, P4, and P5. Step 3.1.4: Splice the features of the same feature map at different scales together through the SPPF module.

4. The foreign object detection method for transmission lines based on deep learning according to claim 3, characterized in that Step 3.2 specifically includes the following steps: Step 3.2.1: Perform convolution operations on the feature maps P3, P4, and P5 extracted by the Backbone network to unify the number of channels of the feature maps. Step 3.2.2: Upsample the feature map P5, fuse it with P4 using the weight fusion method, and obtain the feature map P54 at this time through the node VoVGSCSP. Step 3.2.3: Upsample the feature map P54, fuse it with P3 using the weight fusion method, and obtain the feature map P*3 at this time through the node VoVGSCSP. Thus, the weighted bidirectional feature pyramid BiFPN has completed the top-down feature fusion work. Step 3.2.4: Downsample the feature map P2, fuse it with the feature maps P3 and P54 through the node VoVGSCSP to obtain the feature map P23 at this time. Step 3.2.5: Downsample the feature map P23, fuse it with the feature maps P4 and P54 through the node VoVGSCSP to obtain the feature map P34 at this time. Step 3.2.6: Downsample the feature map P34, fuse it with the feature map P5 through the node VoVGSCSP to obtain the feature map P*5. Thus, BiFPN has completed the bottom-up feature fusion work. Step 3.2.7: Through Steps 3.2.1 - 3.2.6, a feature map that combines high-level semantic information and low-level position information is obtained.

5. The foreign object detection method for transmission lines based on deep learning according to claim 4, characterized in that, Step 3.3 specifically includes the following steps: Before the output of the Head network, introduce the object detection head Dynamic head based on the attention mechanism. Apply the attention mechanism from three different perspectives: scale perception, spatial perception, and task perception, to improve the expression ability of the object detection head of the TL_Yolo model without increasing the computational cost. Dynamic head is the superposition of three kinds of attention, and the stacking of the three kinds of attention is called a DyHead block, where π L , π S , π C represent scale perception attention, spatial perception attention, and task perception attention respectively. The specific composition of Dynamic head is as follows: The scale perception module consists of average pooling + 1x1 convolution + ReLU + sigmoid functions. The spatial perception module consists of index + 3x3 convolution + sigmoid + offset; among which, the index operation uses deformable convolutions of different sizes for feature maps of different sizes; offset is a standard 3x3 convolution; The task perception module consists of average pooling + fully connected layer + ReLU + fully connected layer + Normalize; among which, Normalize is implemented through sigmoid, and the specific formula is: The combination of max and min operations ensures that the output value of sigmoid is in the range of 0 to 1, controls the ON / OFF of the switches of different feature map channels according to different tasks, so as to achieve task perception; the feature maps P3, P4, and P5 after feature fusion are input into the DyHead block to perform scale perception, spatial perception and task perception operations in sequence, and then the feature maps are input into the Decoupled-Head for prediction.

6. The foreign object detection method for transmission lines based on deep learning according to claim 5, characterized in that Step 6 specifically includes the following steps: The Loss calculation of the TL_Yolo model includes two parts. First, calculate the VFL Loss for the classification branch, and its calculation formula is: where p is the predicted value, q is the true label value, and η and γ are weight parameters; for the regression branch, calculate the DistributionFocal Loss and CIoU Loss. Among them, the Distribution Focal Loss uses the cross-entropy function to optimize the probabilities of the two positions on the left and right near the label, so that the model can quickly focus on the values near the label, and its calculation formula is DFL(S i ,S i+1 ) = -((y i+1 - y) log(S i ) + (y - y i ) log(S i+1 )) (4) Among which S i is the sigmod output of the network, i and i + 1 are the interval order, and y is the label value; the CIoU loss L CIOU The calculation formula is as follows: where α is the weight function, ν is used to measure the aspect ratio consistency, b represents the predicted bounding box, and b gt represents the ground truth bounding box, ρ 2 (b, b gt ) represents the Euclidean distance between the centers of the predicted bounding box and the ground truth bounding box, c represents the diagonal distance of the smallest closed region that can contain both the predicted bounding box and the ground truth bounding box, and IOU is the intersection over union of the predicted bounding box and the ground truth bounding box, that is