Power transmission line foreign matter detection method based on improved YOLOv8

By improving the YOLOv8 model, adding variable convolution and hybrid attention mechanisms, and combining context feature enhancement modules and feature processing networks, the problem of identifying complex edge foreign objects and improving feature extraction capabilities in foreign matter detection in transmission lines is solved, achieving higher recognition accuracy and faster model convergence.

CN120047787APending Publication Date: 2025-05-27NANJING ZHENGTU INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411888485.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The current transmission line foreign object detection algorithm has shortcomings in identifying complex edge foreign objects and improving model feature extraction capabilities, resulting in insufficient recognition accuracy and real-timeness to meet actual work requirements.

Method used

By improving the network structure of the YOLOv8 model, adding variable convolution and hybrid attention mechanisms, and building a detection network model containing context feature enhancement modules and feature processing networks, it is trained in combination with redefined model loss functions and transfer learning.

Benefits of technology

The detection network's accurate recognition ability of complex edge foreign objects is improved, feature extraction ability is enhanced, and the problem of difficulty in model training and low recognition accuracy is solved, while accelerating the convergence speed of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047787A_ABST
    Figure CN120047787A_ABST
Patent Text Reader

Abstract

The invention discloses an improved YOLOv8-based power transmission line foreign matter detection method, which comprises the steps of constructing a power transmission line foreign matter detection network model, improving an original framework of a YOLOv8 model in a backbone network and a feature processing network of the model, and fusing deformable convolution, a mixed attention mechanism and a context feature enhancement module. And training the model in combination with a redefined model loss function and transfer learning to obtain a trained foreign matter detection network model, and performing target detection on the unprocessed foreign matter of the power transmission line according to the trained foreign matter detection network model. According to the method, the problems of difficulty in model training, to-be-improved model precision and low target identification accuracy in the field of transmission line foreign matter detection are solved, the convergence speed of the model is accelerated, and the model identification quality is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a foreign object detection method for transmission lines based on improved YOLOv8, belonging to the technical field of object detection. Background Art

[0002] Transmission lines are long-term outdoors, facing a complex set of uncertainties. Any minor problem may trigger a safety accident or even lead to the paralysis of the entire regional power system. For example, branches and wires falling on the lines may cause short circuits, and bird droppings can easily cause insulator flashovers. This requires planned safety inspections of high-voltage transmission lines and their attached equipment to promptly eliminate overhead line faults and avoid serious accidents and losses. Common foreign objects invading transmission lines include bird nests, kites, and hanging balloons. The inspection of transmission lines is an important means to eliminate foreign object intrusions and ensure the stable operation of the lines, so high requirements are put forward for the inspection quality.

[0003] Unmanned aerial vehicle (UAV) autonomous inspection can automatically plan flight paths and collect data, covering professional technologies such as hardware design, computer vision, and communication engineering. On the premise of ensuring accurate and intelligent path planning and high-efficiency data transmission, the accuracy of the image recognition algorithm is crucial. Through the unremitting efforts of researchers, the deep learning algorithms equipped on UAVs have basically met the requirements of actual inspections. The current research focus is on how to improve and optimize to further enhance the recognition accuracy.

[0004] He Jun et al. proposed a detection method based on the deep convolutional neural network EF-YOLO to identify birds on transmission lines. By referring to the feature extraction part of the EfficientNet-lite lightweight network, the EF-YOLO model was proposed, and an appropriate loss function was selected to make it have good detection accuracy and real-time performance. Zou Huijun et al. proposed a method for detecting small target foreign objects on transmission lines based on the improved YOLOv5. By optimizing the original YOLOv5 network structure, the original FPN structure was replaced with BiFPN, reducing the computational amount, and the dataset was expanded by means of scene enhancement and adding noise, making the network more conducive to detecting potential safety hazards caused by small foreign objects. Yang Jianfeng et al. proposed an intrusion recognition method for foreign objects on transmission lines based on the improved Dense-YOLOv3 network model by optimizing the YOLOv3 network. The conditional generative adversarial network algorithm was used to expand the image data containing target foreign objects, solving the problem of fewer foreign object image data samples, and the DenseNet network was introduced, making the network have good recognition effects. Chen Jiachen et al. proposed a method for identifying transmission line defects based on the YOLOv3 network. The spatial pyramid pooling module was introduced into the original YOLOv3 structure, improving the accuracy of detecting images of different sizes, and reducing the network channels of the new model to further lightweight it, reducing the hardware requirements for the server, not only reducing the iteration time, but also improving the overall performance. Since most of the transmission line images obtained by drones have complex backgrounds and the shapes of foreign objects are variable, the accuracy and real-time performance of many current recognition algorithms cannot meet the actual working requirements. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method for detecting foreign objects on transmission lines based on the improved YOLOv8, which improves the detection accuracy by optimizing the network structure of the original YOLOv8 model, making it more suitable for the transmission line inspection task and capable of accurately detecting strange objects.

[0006] The present invention adopts the following technical solutions to solve the above technical problems:

[0007] A method for detecting foreign objects on transmission lines based on the improved YOLOv8 includes the following steps:

[0008] Step 1, obtain a dataset of foreign object images on transmission lines, perform data augmentation on the dataset to obtain an augmented dataset, and divide the augmented dataset into a training set and a validation set;

[0009] Step 2: Construct a foreign object detection network model based on the improved YOLOv8, including an input end, a backbone network, a feature processing network, and a detection head. Among them, the backbone network includes the first to fifth basic convolution modules, the first to fourth hybrid attention fusion modules, a context feature enhancement module, and a spatial pyramid pooling module combined with deformable convolution. The first basic convolution module, the second basic convolution module, the first hybrid attention fusion module, the third basic convolution module, the second hybrid attention fusion module, the fourth basic convolution module, the third hybrid attention fusion module, the fifth basic convolution module, the fourth hybrid attention fusion module, the context feature enhancement module, and the spatial pyramid pooling module combined with deformable convolution are connected in sequence. The input end is connected to the first basic convolution module, and the outputs of the second hybrid attention fusion module, the third hybrid attention fusion module, and the spatial pyramid pooling module combined with deformable convolution are used as the first, second, and third outputs of the backbone network respectively.

[0010] Step 3: Select a classification loss function and a regression loss function, and use the training set and the validation set to train the foreign object detection network model constructed in Step 2 to obtain a trained foreign object detection network model.

[0011] Step 4: Use the trained foreign object detection network model to detect the foreign object images of the transmission line collected in real time to obtain the foreign object detection results.

[0012] As a preferred solution of the present invention, the feature processing network includes the sixth to seventh basic convolution modules and the fifth to eighth hybrid attention fusion modules. The third output of the backbone network is fused with the second output of the backbone network after an upsampling operation to obtain a first fusion result. The first fusion result is fused with the first output of the backbone network after passing through the fifth hybrid attention fusion module and an upsampling operation in sequence to obtain a second fusion result. The second fusion result is fused with the output of the fifth hybrid attention fusion module after passing through the sixth hybrid attention fusion module and the sixth basic convolution module in sequence to obtain a third fusion result. The third fusion result is fused with the third output of the backbone network after passing through the seventh hybrid attention fusion module and the seventh basic convolution module in sequence to obtain a fourth fusion result. The fourth fusion result passes through the eighth hybrid attention fusion module to obtain the third output of the feature processing network, and the outputs of the sixth and seventh hybrid attention fusion modules are used as the first and second outputs of the feature processing network respectively.

[0013] As a preferred solution of the present invention, the detection head includes decoupled detection heads of large, medium, and small scales. Each scale of the decoupled detection head is divided into two branches, and each branch includes two basic convolution modules and a two-dimensional convolution layer connected in sequence. The first, second, and third outputs of the feature processing network are respectively connected to the decoupled detection heads of large, medium, and small scales.

[0014] As a preferred embodiment of the present invention, the structures of the basic convolution modules are the same, each including a convolution layer, batch normalization, and a Silu activation function connected in sequence.

[0015] As a preferred embodiment of the present invention, the structures of the first to eighth hybrid attention fusion modules are the same. By combining the convolutional structure for fast feature embedding, the spatial attention mechanism, and the channel attention mechanism, a hybrid attention fusion module is obtained. The spatial attention mechanism includes a max pooling layer, the first to third 3×3 convolutional layers, and a first convolutional layer. The channel attention mechanism includes a global pooling layer, a fourth 3×3 convolutional layer, and the first to second fully connected layers.

[0016] The input of the hybrid attention fusion module passes through the convolutional structure for fast feature embedding to obtain the output of the convolutional structure for fast feature embedding, and the output of the convolutional structure for fast feature embedding is used as the input of the spatial attention mechanism. The output of the convolutional structure for fast feature embedding passes through the first 3×3 convolutional layer, the second 3×3 convolutional layer, the first convolutional layer, the max pooling layer, the third 3×3 convolutional layer, and the first Sigmoid activation function in sequence to obtain the output of the first Sigmoid activation function. The result of multiplying the output of the first Sigmoid activation function by the output of the first 3×3 convolutional layer is used as the input of the channel attention mechanism.

[0017] The input of the channel attention mechanism passes through the fourth 3×3 convolutional layer, the global pooling layer, the first fully connected layer, the ReLU activation function, the second fully connected layer, and the second Sigmoid activation function in sequence to obtain the output of the second Sigmoid activation function. The result of multiplying the output of the second Sigmoid activation function by the output of the fourth 3×3 convolutional layer is used as the output of the hybrid attention fusion module.

[0018] As a preferred embodiment of the present invention, the spatial pyramid pooling module combined with deformable convolution includes a first deformable convolutional layer, a basic convolution module combined with deformable convolution, and the first to third max pooling layers. The input of the spatial pyramid pooling module combined with deformable convolution passes through the first deformable convolutional layer, batch normalization, the SiLU activation function, and the first max pooling layer in sequence to obtain the output of the first max pooling layer. The output of the first max pooling layer passes through the second max pooling layer and the third max pooling layer to obtain the outputs of the second pooling layer and the third pooling layer respectively. After fusing the output of the SiLU activation function, the output of the first max pooling layer, the output of the second max pooling layer, and the output of the third max pooling layer, it is used as the input of the basic convolution module combined with deformable convolution, and the output of the basic convolution module combined with deformable convolution is used as the output of the spatial pyramid pooling module combined with deformable convolution.

[0019] The basic convolution module combined with variable convolution includes a second variable convolution layer, batch normalization, and a SiLU activation function connected in sequence.

[0020] As a preferred solution of the present invention, the context feature enhancement module includes first to third dilated convolution layers, and the dilation rates of the first to third dilated convolution layers are 1, 3, and 5 respectively. The input of the context feature enhancement module is respectively passed through the first, second, and third dilated convolution layers, and the outputs of the first, second, and third dilated convolution layers are fused as the output of the context feature enhancement module.

[0021] As a preferred solution of the present invention, in step 3, the expression of the classification loss function is as follows:

[0022]

[0023] Among them, VFL(p,q) represents the classification loss, p is the confidence predicted by the model, q is the true confidence score represented by the label, and P γ is the focal coefficient, and γ is the adjustment parameter;

[0024] The regression loss function combines the Wise-IoU loss function and the distribution focal loss function. Among them, the expression of the Wise-IoU loss function is as follows:

[0025]

[0026] L IoU = 1 - R IoU

[0027] Among them, L WIoU represents the Wise-IoU loss function, α is a hyperparameter, δ is an adjustment parameter, (x,y) and (x gt ,y gt ) are the coordinates of the centers of the anchor box and the target box respectively, W g and H g are the width and height of the minimum bounding box respectively, and * represents the operation of separating W g and H g , R IoU represents the intersection over union of the predicted box and the true box, and β is the outlier parameter. The formula is:

[0028]

[0029] Among them, is the gradient gain of the monotonic focusing coefficient, is the sliding average of the momentum m, t is the epoch value of the complete dataset after one training, and n is the batch size;

[0030] The expression of the distribution focal loss function is as follows:

[0031] DFL(S i ,S i+1 ) = -((y i+1 -y)log(S i )+(y-y i )log(S i+1 ))

[0032] Among them, DFL(S i ,S i+1 ) represents the distribution focal loss function. i and i + 1 are the position indices after discretization of the predicted bounding boxes respectively. S i and S i+1 are the probability values of the i-th and (i + 1)-th predicted bounding boxes output by the model respectively. y i and y i+1 are the actual coordinate values of the i-th and (i + 1)-th predicted bounding boxes respectively. y is the true continuous coordinate value of the target position.

[0033] A computer device includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, the steps of the foreign object detection method for transmission lines based on the improved YOLOv8 are implemented.

[0034] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the foreign object detection method for transmission lines based on the improved YOLOv8 are implemented.

[0035] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:

[0036] 1. By improving the traditional YOLOv8 model and adding deformable convolution and hybrid attention mechanism, the detection network can accurately identify foreign objects with complex edges, and at the same time improve the model's ability to extract features.

[0037] 2. The network model constructed by the present invention integrates a context feature enhancement module, and combines a redefined model loss function and transfer learning to train the model, solving the problems of difficult model training, low model accuracy, and low recognition accuracy in the field of foreign object recognition for transmission lines. At the same time, it accelerates the convergence speed of the model and ensures the recognition quality of the model. Description of the Drawings

[0038] Figure 1 is the structural diagram of the foreign object detection network model based on the improved YOLOv8 constructed by the present invention;

[0039] Figure 2 It is the structural diagram of the hybrid attention fusion module (C2F-Ab) of the present invention;

[0040] Figure 3 It is the structural diagram of the deformable convolutional layer (DCN) of the present invention;

[0041] Figure 4 It is the structural diagram of the spatial pyramid pooling module combined with deformable convolution (SPPF-Dcn) and the basic convolutional module combined with deformable convolution (CBS-Dcn) of the present invention;

[0042] Figure 5 It is the structural diagram of the context feature enhancement module (FEM) of the present invention. Specific embodiments

[0043] The following details the embodiments of the present invention, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation of the present invention.

[0044] The present invention proposes a foreign object detection method for transmission lines based on improved YOLOv8, including the following steps:

[0045] Step 1: Obtain transmission line images. Due to the particularity of transmission line images, it is difficult to obtain data samples, and the sample quality is generally low. Therefore, data enhancement operations are performed on the transmission line images, and common image enhancement methods in power object detection are used, such as flipping, rotating, scaling, cropping, adding Gaussian noise, and brightness adjustment. The enhanced image set is divided into a training set and a validation set.

[0046] Step 2: Construct a foreign object detection network model. As Figure 1 shown, the model is divided into four parts, namely the input end (Input), the backbone network (Backbone), the feature processing network (Neck), and the detection head (Head). Among them, the input end is the foreign object image of the transmission line after data enhancement. The backbone network includes a basic convolutional module (CBS), a hybrid attention fusion module (C2F-Ab), a context feature enhancement module (FEM), and a spatial pyramid pooling module combined with deformable convolution (SPPF-Dcn). The feature processing network adopts a network architecture that combines a feature pyramid structure and a feature aggregation network to extract information of different dimensions during the upsampling and downsampling processes of the features. The detection head adopts decoupled detection heads of large, medium, and small scales, and each decoupled detection head is divided into two independent branches of classification and regression to improve the detection accuracy through two convolutional channels.

[0047] In the backbone network (Backbone), the improved C2F module is combined with the hybrid attention mechanism to perform weighted feature extraction and enhance the extraction capability of the main module, namely C2F-Ab. At the same time, the variable convolution is combined with the CBS module to design the basic convolution module combined with variable convolution (CBS-Dcn) and the spatial pyramid pooling module combined with variable convolution (SPPF-Dcn), so that the network also has a good modeling effect on irregular objects in the transmission line.

[0048] The image is input into the backbone network for processing. First, it passes through several CBS modules in the backbone network, including convolution (conv), batch normalization (BN), and activation function SiLU. Then it passes through the C2F-Ab module, in which the bottleneck layers originally connected in series are connected using the idea of ​​gradient diversion, which optimizes the module structure and effectively avoids the deterioration of convergence when the depth is too deep. While optimizing the network structure to achieve lightweight, rich gradient flow information can be obtained, which improves the overall performance of the YOLOv8 model. Then the output feature value is combined with the hybrid attention mechanism, such as Figure 2 As shown in the figure. The spatial attention mechanism uses a bottleneck layer structure, and performs lightweight operations on the channel through a small-size convolution kernel combined with convolution. First, the input is dimensionalized by a 3×3 convolution, so that the maximum pooling layer (MaxPooling) can filter data more effectively and intuitively. Then, the weight is scaled to between 0 and 1 through a Sigmoid activation function, and multiplied point by point with the original feature map to obtain a feature map containing the weight. The channel attention mechanism reduces the dimension of the input vector X by global pooling and converts it into an output of 1×1×1×C. The fully connected layer (FC) layer uses a cheap convolution C-Conv with kernel_size=1. The cheap convolution Cheap Conv can ensure the effect while ensuring fewer parameters. The channel is filtered by the activation function ReLU, and data enhancement is performed again through a layer of FC. Finally, the weight is scaled to the interval [0,1] by a Sigmoid function, and then multiplied point by point with the original feature map.

[0049] After passing through several CBS and C2F-Ab modules, the features are sent to the contextual feature enhancement module (FEM), which increases the detection receptive field through dilated convolution with dilation rates of 1, 3, and 5. At the same time, the feature information of the target to be detected and its surroundings is combined to enhance the feature information of small targets in disguise, thereby enhancing the algorithm's ability to understand the target. The FEM structure is as follows: Figure 5As shown in the figure, the features are processed by dilated convolutions with dilation rates of 1, 3, and 5, taking into account both the features of the target itself and the relative features of the target's surrounding environment. Then, the three processed feature maps are concatenated by Concat and input into the next layer of the network structure for processing.

[0050] As Figure 4 shown, multi-scale feature extraction and fusion are performed in the Spatial Pyramid Pooling Module with Deformable Convolution (SPPF-Dcn). First, in this module, the deformable convolution (DCN) is introduced into the CBS module to propose a basic convolution module with deformable convolution (CBS-Dcn). The deformable convolution can automatically adapt and adjust according to the shape and proportion of the object to be measured. By using an irregular convolution kernel, it solves the drawback of insufficient sampling in the traditional fixed rectangular structure, enabling the network model to better simulate the deformation of objects. Its principle is as Figure 3 shown, the feature output by the deformable convolution is:

[0051]

[0052] In the formula, R ∈ {(-1, -1), (-1, 0),..., (0, 1), (1, 1)}, ω(p n ) represents the convolution kernel weight value at the position of p n , represents the feature map to be detected, p n represents the position element in R, and Δp n represents the offset added to the regular convolution sampling point. The pixel at the sampling point position with the offset is calculated using the bilinear interpolation method, and the formula is as follows:

[0053]

[0054] Where p = p 0 + p n + Δp n represents any position in the region, q represents all spatial positions in the input feature map, that is, the four integer points around p. represents the value at all integer positions in the feature map, and G(q, p) is a bilinear interpolation kernel function of a two-dimensional kernel, which can be decomposed into two one-dimensional kernels, and the formula is as follows:

[0055] G(q, p) = g(q x , p x ) · g(q y , p y )

[0056] In the formula, g(a, b) = max{0, 1 - |a - b|}.

[0057] After the backbone network extracts multi-layer eigenvalues from the original image, the feature map will enter the feature processing network (Neck). Here, the path aggregation network is combined with the feature pyramid structure, and different levels of features are aggregated through upsampling and downsampling operations. The top-down part mainly realizes the fusion of different levels of features by upsampling and fusing with coarser-grained feature maps, while the bottom-up part fuses feature maps from different levels by using a convolutional layer.

[0058] The specific operations are as follows. The top-down part realizes the fusion of different levels of features by upsampling and fusing with coarser-grained feature maps, which is mainly divided into the following steps:

[0059] a. Upsample the last layer of the feature map to obtain a finer feature map;

[0060] b. Fuse the upsampled feature map with the previous layer of the feature map to obtain a richer feature representation;

[0061] c. Repeat the above two steps until the highest layer is reached.

[0062] The bottom-up part mainly fuses feature maps from different levels by using a convolutional layer, which is mainly divided into the following steps:

[0063] d. Convolve the bottom layer of the feature map to obtain a richer feature representation;

[0064] e. Fuse the convolved feature map with the previous layer of the feature map to obtain a richer feature representation;

[0065] f. Repeat the above two steps until the highest layer is reached.

[0066] The detection head (Head) is used to perform object detection on the feature pyramid. Through two branches, it passes through two CBS modules and a conventional two-dimensional convolution to detect objects. The decoupled detection head using this method has the effect of accelerating convergence and improving accuracy.

[0067] Step 3: Select the loss function of the network. The classification loss function selects VFI (Varifocal Loss), and the regression loss function adopts the optimized Wise-IoU Loss + Distribution Focal Loss (DFL).

[0068] Varifocal Loss function: To address the problem of uneven positive and negative samples, the Varifocal Loss function is used to propose an asymmetric weighted operation. The formula of the Varifocal Loss function is as follows:

[0069]

[0070] Among them, p is the confidence (probability) predicted by the model (IoU-aware classification score, IACS), which combines the confidence of classification and the IoU information of the bounding box, and is the predicted probability value. q is the label representing the true confidence score, P γ is a focus coefficient, used to reduce the attention to easy-to-separate samples and increase the attention to difficult-to-separate samples. γ is a tuning parameter, used to balance the loss weights of positive and negative samples. Varifocal Loss no longer processes positive and negative samples symmetrically. By considering their different importance levels, it highlights positive samples as the main body.

[0071] Wise-IoU Loss function: IoU is the intersection over union of the predicted box and the ground truth box, which is an indicator to measure the accuracy of object detection. The larger the IoU, the higher the overlap degree between the predicted box and the ground truth box. Compared with IoU, CIoU is more sensitive to the object size, but CIoU calculates the loss of the bounding box. It adds the loss calculation of the aspect ratio, but does not consider the balance problem of the dataset samples themselves. Most of the previous loss functions rarely consider the quality of the labeled examples in the dataset itself, but continuously strive to strengthen the fitting ability of the bounding box loss, resulting in a great impact of some low-quality annotations on the detection performance. Therefore, in the present invention, using Wise-IoU instead of CIoU can improve the accuracy of small object detection. The calculation formula is as follows:

[0072]

[0073] L IoU = 1 - R IoU

[0074] where (x, y) and (x gt , y gt ) are the coordinates of the center points of the anchor box and the target box respectively; W g and H g are the sizes of the smallest enclosing box; * represents the operation of separating W g and H g ; R IoU represents the intersection over union, which is a common quantity to measure the overlapping degree of the predicted box and the ground truth box. α is a hyperparameter, and δ is a tuning parameter to balance the loss weight. α and δ are set to 1.9 and 3.0 respectively.

[0075] By introducing an outlier parameter β to describe the quality of the anchor box, β is negatively correlated with the anchor box quality. The outlier calculation formula is:

[0076]

[0077] Among them, is the gradient gain of the monotonic focusing coefficient, which is the same as L IoU defined as the same, where * indicates that it will be continuously calculated and changed according to the situation of each object detection during the training process; is the moving average value of momentum m, and introducing can dynamically adjust the highest gradient gain according to the training process. The calculation formula of momentum m is:

[0078]

[0079] Among them, t is the epoch value of a complete dataset after one training, and n is the value of the batch size. The significance of introducing momentum m is that after t rounds of training, the Wise-IoU Loss assigns small gradient gains to low-quality anchor boxes to reduce harmful gradients.

[0080] Distribution Focal Loss function: When detecting the input image, there are often multiple object occlusion or overlap situations. If there is an overlap situation, at this time, neither the annotation box nor the detection box can truly reflect the image semantics. Therefore, a more accurate bounding box representation method is needed at this time, changing the Dirac distribution representation method to a more general probability representation method, and the formula is as follows:

[0081]

[0082] is the predicted output of the network. This formula indicates that the predicted coordinates are not only achieved by adjusting the number of channels through convolution, but rather the network predicts a probability distribution and weights the coordinates based on this probability distribution to obtain the predicted coordinates.

[0083] Since the network cannot directly output a continuous probability distribution and uses discrete probability points to represent, the formula is as follows:

[0084]

[0085] However, if directly using this for training, it will make the representation space of the probability too large, making it difficult for the network to optimize and converge. Because the interval points near the true coordinates have a higher probability, the Distribution Focal Loss, abbreviated as DFL, is proposed, and the formula is as follows:

[0086] DFL(S i , S i+1 ) = -((y i+1 - y) log(S i ) + (y - y i ) log(S i++ ))

[0087] S i and S i+1 are the probability values output by the model. These probability values represent the confidence scores at a certain discrete position. y i and y i++ are the actual coordinate values of the discrete positions. y is the true continuous coordinate value of the target position, and i represents the index value of the predicted bounding box. It can be seen from the formula that when y i+1 is very close to y i and the S probability output is large, the DFL is small at this time, making the distribution approach the annotation center. Therefore, DFL can enable the newly proposed network to focus on the values near the target y faster, increase their probabilities, and accelerate convergence.

[0088] Step 4: Train the object detection network model constructed in Step 1 according to the loss function defined in Step 2. Introduce transfer learning during the training process to obtain a trained object detection network model, and use the trained object detection network model to detect foreign objects in the unprocessed transmission line images.

[0089] Embodiment

[0090] Step 1: Use the self-made dataset as the training set and train all networks on NVIDIA - RTX2080ti.

[0091] Step 2: During the training, for each batch, the batch size is set to 8, the IoU threshold is set to 0.5, and the number of iterations is 100. The setting ratio of the training set to the validation set for each type of foreign object on the transmission line is 9:1, and the quantity ratio between the two and the test set is also set to 9:1. Use the Adam optimizer with β = 0.9 to achieve optimization, and the initial learning rate is 10e -5 .

[0092] Step 3: Use the self-made test set or brand-new drone-shot line images to test the model. The selected test images are not processed by enhancement. Since it is difficult to obtain the image data of each foreign object on the transmission line, the original sample size of this experiment did not reach the expectation. Therefore, the model is pre-trained on the COCO dataset used as the source domain, the backbone part of the network is frozen, and the unfrozen part is fine-tuned. Then unfreeze it, change the parameters of the feature extraction network, and conduct a complete training on the network model.

[0093] The present invention uses five evaluation metrics to measure the accuracy of the predicted trajectory, ensuring the dimensional diversity of the evaluation metrics.

[0094] (1) Precision

[0095] Precision is the accuracy rate, abbreviated as P, which is used to measure the precision of the model. The accuracy rate originally referred to the proportion of the correctly predicted part in all the prediction results. In the present invention, it specifically represents the proportion of the number of correctly predicted foreign objects in the total number of foreign objects predicted by the model. The calculation formula is as shown in the figure:

[0096]

[0097] (2) Recall

[0098] Recall is the recall rate, abbreviated as R. It refers to the proportion of the number of correctly predicted positive samples in the total number of positive samples. In the present invention, it refers to the proportion of the number of correctly predicted foreign objects in the total number of foreign objects. The calculation formula is as follows:

[0099]

[0100] In the formula, TP (True positives) refers to the number of foreign objects correctly identified; FP (False positives) refers to the number of non-foreign objects identified as foreign objects; FN (Fales negatives) refers to the number of foreign objects identified as non-foreign objects.

[0101] (3) F1 Score

[0102] The F1 score, also known as the balanced score, is the harmonic mean of the precision rate and the recall rate. The calculation formula is as follows:

[0103]

[0104] (4) Average Precision (AP)

[0105] Using only the precision rate and the recall rate to evaluate the model is too one-sided. It is necessary to use a comprehensive index AP (Average Precision) to measure the detection performance of the model. In the present invention, it refers to the detection precision of various foreign objects. The calculation formula of the AP value is as follows:

[0106] AP = ∫ 0 1 P(R)dr

[0107] (5) mean Average Precision (mAP)

[0108] The mean Average Precision, abbreviated as mAP, is the mean of the detection target precision value AP. Since the precision rate and the recall rate are two contradictory indicators, it is necessary to comprehensively consider the two using mAP. Briefly speaking, mAP is a comprehensive index to measure the detection effect of the algorithm. In this article, it refers to the average detection precision of all types of overhead line foreign objects. The calculation formula is as follows:

[0109]

[0110] The larger the mAP value is, the better the comprehensive performance of the algorithm is.

[0111] Based on the same inventive concept, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the aforementioned foreign object detection method for transmission lines based on improved YOLOv8 are implemented.

[0112] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the aforementioned foreign object detection method for transmission lines based on improved YOLOv8 are implemented.

[0113] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0114] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0115] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the process Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for implementing the functions specified in one block or a plurality of blocks.

[0117] The above embodiments are only used to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the present invention.

Claims

1. A method for detecting foreign objects in a power transmission line based on improved YOLOv8, characterized in that: The steps include: Step 1, obtaining a data set of foreign body images of power transmission lines, and performing data enhancement on the data set to obtain an enhanced data set, and dividing the enhanced data set into a training set and a verification set; Step 2, constructing a foreign body detection network model based on improved YOLOv8, including an input end, a backbone network, a feature processing network and a detection head; wherein the backbone network includes the first to fifth basic convolution modules, the first to fourth hybrid attention fusion modules, a context feature enhancement module and a spatial pyramid pooling module combined with variable convolution, and the first basic convolution module, the second basic convolution module, the first hybrid attention fusion module, the third basic convolution module, the second hybrid attention fusion module, the fourth basic convolution module, the third hybrid attention fusion module, the fifth basic convolution module, the fourth hybrid attention fusion module, the context feature enhancement module and the spatial pyramid pooling module combined with variable convolution are connected in sequence; the input end is connected to the first basic convolution module, and the outputs of the second hybrid attention fusion module, the third hybrid attention fusion module and the spatial pyramid pooling module combined with variable convolution are respectively used as the first, second and third outputs of the backbone network; Step 3, selecting a classification loss function and a regression loss function, and using the training set and the validation set to train the foreign body detection network model constructed in step 2 to obtain a trained foreign body detection network model; Step 4: Use the trained foreign object detection network model to detect the foreign object images of the transmission line collected in real time to obtain the foreign object detection results.

2. According to claim 1, the method for detecting foreign objects in a power transmission line based on improved YOLOv8 is characterized in that: The feature processing network includes sixth to seventh basic convolution modules and fifth to eighth hybrid attention fusion modules, and the third output of the backbone network is fused with the second output of the backbone network after upsampling operation to obtain a first fusion result; The first fusion result is sequentially passed through the fifth hybrid attention fusion module and the upsampling operation, and then fused with the first output of the backbone network to obtain the second fusion result; The second fusion result is sequentially passed through the sixth hybrid attention fusion module and the sixth basic convolution module, and then fused with the output of the fifth hybrid attention fusion module to obtain a third fusion result; The third fusion result is sequentially passed through the seventh hybrid attention fusion module and the seventh basic convolution module, and then fused with the third output of the backbone network to obtain a fourth fusion result; The fourth fusion result is passed through the eighth mixed attention fusion module to obtain the third output of the feature processing network, and the outputs of the sixth and seventh mixed attention fusion modules are used as the first and second outputs of the feature processing network respectively.

3. The method for detecting foreign objects in a power transmission line based on improved YOLOv8 according to claim 2, characterized in that: The detection head includes decoupled detection heads of three scales: large, medium and small. The decoupled detection head of each scale is divided into two branches, each branch includes two basic convolution modules connected in sequence and a two-dimensional convolution layer; the first, second and third outputs of the feature processing network correspond to the decoupled detection heads of large, medium and small scales, respectively.

4. The method for detecting foreign objects in a power transmission line based on improved YOLOv8 according to claim 3 is characterized in that: The structures of the basic convolution modules are the same, including sequentially connected convolutional layers, batch normalization, and Silu activation functions.

5. The method for detecting foreign objects in a power transmission line based on improved YOLOv8 according to claim 2, characterized in that: The structures of the first to eighth hybrid attention fusion modules are the same, which combine the convolution structure of fast feature embedding, the spatial attention mechanism and the channel attention mechanism to obtain a hybrid attention fusion module; the spatial attention mechanism includes the maximum pooling layer, the first to third 3×3 convolution layers and the first convolution layer, and the channel attention mechanism includes the global pooling layer, the fourth 3×3 convolution layer and the first to second fully connected layers; The input of the hybrid attention fusion module is passed through the convolution structure of fast feature embedding to obtain the output of the convolution structure of fast feature embedding, and the output of the convolution structure of fast feature embedding is used as the input of the spatial attention mechanism; The output of the convolution structure of fast feature embedding is sequentially passed through the first 3×3 convolution layer, the second 3×3 convolution layer, the first convolution layer, the maximum pooling layer, the third 3×3 convolution layer and the first Sigmoid activation function to obtain the output of the first Sigmoid activation function. The output of the first Sigmoid activation function is multiplied by the output of the first 3×3 convolution layer as the input of the channel attention mechanism. The input of the channel attention mechanism passes through the fourth 3×3 convolution layer, the global pooling layer, the first fully connected layer, the ReLU activation function, the second fully connected layer and the second Sigmoid activation function in sequence to obtain the output of the second Sigmoid activation function. The output of the second Sigmoid activation function is multiplied by the output of the fourth 3×3 convolution layer as the output of the hybrid attention fusion module.

6. The method for detecting foreign objects in a power transmission line based on improved YOLOv8 according to claim 1, characterized in that: The spatial pyramid pooling module combined with variable convolution includes a first variable convolution layer, a basic convolution module combined with variable convolution, and first to third maximum pooling layers; the input of the spatial pyramid pooling module combined with variable convolution is sequentially subjected to the first variable convolution layer, batch normalization, SiLU activation function and the first maximum pooling layer to obtain the output of the first maximum pooling layer, the output of the first maximum pooling layer is sequentially subjected to the second maximum pooling layer and the third maximum pooling layer to obtain the output of the second pooling layer and the third pooling layer respectively, the output of the SiLU activation function, the output of the first maximum pooling layer, the output of the second maximum pooling layer and the output of the third maximum pooling layer are fused as the input of the basic convolution module combined with variable convolution, and the output of the basic convolution module combined with variable convolution is used as the output of the spatial pyramid pooling module combined with variable convolution; The basic convolution module combined with variable convolution includes a second variable convolution layer, batch normalization and SiLU activation function connected in sequence.

7. The method for detecting foreign objects in a power transmission line based on improved YOLOv8 according to claim 1, characterized in that: The context feature enhancement module includes first to third dilated convolutional layers, and the dilation rates of the first to third dilated convolutional layers are 1, 3 and 5 respectively. The input of the context feature enhancement module passes through the first, second and third dilated convolutional layers respectively, and the outputs of the first, second and third dilated convolutional layers are fused as the output of the context feature enhancement module.

8. The method for detecting foreign objects in a power transmission line based on improved YOLOv8 according to claim 1, characterized in that: In step 3, the expression of the classification loss function is as follows: Among them, VFL(p,q) represents the classification loss, p is the confidence of the model prediction, q is the confidence score of the label representation, P γ is the focal coefficient, γ is the adjustment parameter; The regression loss function combines the Wise-IoU loss function and the distribution focus loss function, where The expression of Wise-IoU loss function is as follows: L IoU =1-R IoU Among them, L WIoU represents the Wise-IoU loss function, α is a hyperparameter, δ is a tuning parameter, (x, y) and (x gt ,y gt ) are the coordinates of the center points of the anchor box and the target box respectively, W g and H g are the width and height of the minimum bounding box respectively, * indicates that W g and H g The separated operation, R IoU It represents the intersection-over-union ratio of the predicted box and the true box, β is the outlier parameter, and the formula is: in, is the gradient gain of the monotonic focusing coefficient, is the sliding average of momentum m, t is the epoch value of one training of the complete dataset, and n is the batch size; The expression of the distribution focus loss function is as follows: DFL(S i ,S i+1 )=-((y i+1 -y)log(S i )+(y-y i )log(S i+1 )) Among them, DFL(S i ,S i+1 ) represents the distribution focus loss function, i and i+1 are the position indexes after the discretization of the predicted bounding box, S i and S i+1 are the probability values ​​of the i-th and i+1-th predicted bounding boxes output by the model, y i and i+1 are the actual coordinate values ​​of the i-th and i+1-th predicted bounding boxes respectively, and y is the true continuous coordinate value of the target position.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, the steps of the power transmission line foreign object detection method based on improved YOLOv8 are implemented as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for detecting foreign objects in a power transmission line based on improved YOLOv8 are implemented.

Citation Information

Cited By

  • Railway tunnel foreign matter detection method based on radar point cloud and RGB image fusion

    CN121186807A

  • Railway tunnel foreign matter detection method based on radar point cloud and RGB image fusion

    CN121186807B