Power transmission line foreign matter detection method based on image vision
By introducing the SKAttention module and the bidirectional feature fusion module in the YOLOv8 network model, the feature extraction and detection of transmission line images is solved, and the problems of low efficiency and low accuracy of transmission line foreign matter detection in the prior art are achieved, achieving more efficient and accurate detection effects.
Patent Information
- Application Number
- CN202510372172.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
AI Technical Summary
There is a lack of a reasonably designed foreign object detection method for power line based on image vision in the prior art, and it is difficult to effectively detect foreign objects such as bird's nests and kites in the transmission line image, resulting in low detection efficiency and low accuracy.
The YOLOv8 network model is adopted, and the SKAttention module and the bidirectional feature fusion module are added to perform feature extraction and detection of foreign objects in the transmission line image. The bidirectional feature fusion module fuses multi-level features, while the SKAttention module strengthens key features to improve detection accuracy.
Through this method, the accuracy and efficiency of foreign matter detection in transmission lines can be significantly improved, and the accuracy and reliability of detection results can be ensured.
Smart Images

Figure CN120219361A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of transmission line image processing, and particularly relates to a foreign object detection method for transmission lines based on image vision. Background Art
[0002] As a key component of the power system, the safe operation of transmission lines has a direct impact on the stability of the entire power system. Common foreign objects on transmission lines include bird nests, kites, etc. The intrusion of foreign objects into transmission lines may lead to short circuits, power losses, and even large-scale power outages, seriously affecting social production and daily life. At present, grid inspections rely on manual operations, which are inefficient, inconvenient to operate, and have potential safety hazards. With the development of computer vision and machine learning technologies, advanced technologies have begun to be integrated into the field of transmission line detection. The detection of foreign objects on transmission lines relies on image acquisition devices carried by unmanned aerial vehicles and deep learning algorithms to improve the accuracy and efficiency of detection.
[0003] Therefore, there is currently a lack of a reasonably designed foreign object detection method for transmission lines based on image vision. The foreign objects in the transmission line image are marked by the minimum bounding rectangle, and the SKAttention module and the bidirectional feature fusion module are added based on the YOLOv8 network model to extract and detect the features of the foreign objects in the transmission line image. The bidirectional feature fusion module fuses multi-level features, and the SKAttention module strengthens key features to improve the detection accuracy. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a foreign object detection method for transmission lines based on image vision in view of the above-mentioned deficiencies in the prior art. The method steps are simple and reasonably designed. The foreign objects in the transmission line image are marked by the minimum bounding rectangle, and the SKAttention module and the bidirectional feature fusion module are added based on the YOLOv8 network model to extract and detect the features of the foreign objects in the transmission line image. The bidirectional feature fusion module fuses multi-level features, and the SKAttention module strengthens key features to improve the detection accuracy.
[0005] To solve the above technical problem, the technical solution adopted by the present invention is: A foreign object detection method for transmission lines based on image vision, the method comprising the following steps:
[0006] Step 1: Acquisition and annotation of transmission line images:
[0007] Obtain the power transmission line image, and use a computer to label the foreign objects in the power transmission line image with the minimum bounding rectangle using the labelimg annotation software to obtain the labeled power transmission line image, which is recorded as the training set; among them, a minimum bounding rectangle in the labeled power transmission line image represents a true annotation box, and the marking information of each true annotation box includes the x coordinate, y coordinate, box width, and box height of the center point; among them, the foreign objects are bird nests and kites;
[0008] Step 2: Construct an improved YOLOv8 network model for detecting foreign objects on power transmission lines:
[0009] The improved YOLOv8 network model for detecting foreign objects on power transmission lines includes a YOLOv8 model and an SKAttention module; the YOLOv8 model includes a backbone network, a bidirectional feature fusion module, and a Head module. The backbone network includes a backbone layer, a first-stage layer, a second-stage layer, a third-stage layer, a fourth-stage layer, and an SPPF module. The backbone layer is a CBS module, and the first-stage layer to the fourth-stage layer all include a CBS module and a C2f module; the SKAttention module includes a first SKAttention module, a second SKAttention module, and a third SKAttention module. The output of the C2f module in the second-stage layer is connected to the first SKAttention module, the output of the C2f module in the third-stage layer is connected to the second SKAttention module, and the third SKAttention module is connected between the C2f module in the fourth-stage layer and the SPPF module;
[0010] The bidirectional feature fusion module includes a first upsampling module, a second upsampling module, a first downsampling module, a second downsampling module, a first C2f module, a second C2f module, a third C2f module, a fourth C2f module, a first CBS module, and a second CBS module;
[0011] Step 3: Feature extraction of the power transmission line image:
[0012] Pass the labeled power transmission line image through the improved YOLOv8 network model for detecting foreign objects on power transmission lines to obtain a first feature map, a second feature map, and a third feature map;
[0013] Step 4: Prediction after feature extraction of the power transmission line image:
[0014] The computer processes and detects the first feature map, the second feature map, and the third feature map through the Head module respectively, and obtains the x coordinate X of the center point of the prediction box, the y coordinate Y of the center point of the prediction box, the width w' of the prediction box, and the height h' of the prediction box from the first four channel feature maps; obtains the confidence value of the prediction box from the fifth channel feature map; and obtains the classification probability of the prediction box from the last two channel feature maps.
[0015] Step Five: Training of the improved YOLOv8 network model for overhead transmission line foreign object detection:
[0016] Step 501: Input the x coordinate X of the center point of the prediction box, the y coordinate Y of the center point of the prediction box, the width w' of the prediction box, the height h' of the prediction box, the confidence value, and the classification probability in Step Four. The computer calculates according to L loss = 0.5L box + L conf + 0.1L cls to obtain the total loss function L loss ; where L box represents the optimized bounding box regression loss, L conf represents the confidence loss, and L cls represents the classification loss;
[0017] Step 502: The computer uses the SGD optimizer and utilizes the total loss function L loss to perform iterative training on the improved YOLOv8 network model for overhead transmission line foreign object detection until all the training sets are trained, completing one iteration of training;
[0018] Step 503: Repeat Step 502 for iterative training until the preset number of iterations is satisfied, obtaining the trained improved YOLOv8 network model for overhead transmission line foreign object detection;
[0019] Step Six: Use the trained improved YOLOv8 network model for overhead transmission line foreign object detection to detect foreign objects in subsequent overhead transmission line images.
[0020] For the above method for overhead transmission line foreign object detection based on image vision, further, in Step One, the acquisition of the overhead transmission line image is as follows:
[0021] Step 101: Use a drone equipped with a camera to collect images of the overhead transmission line to be detected, obtain the overhead transmission line image, and transmit it to the computer to obtain the original overhead transmission line image;
[0022] Step 102: Use the computer to perform image processing on the original overhead transmission line image to obtain the enhanced overhead transmission line image; where the image processing includes vertical flipping, horizontal flipping, random brightness adjustment, affine transformation, contrast enhancement, or Gaussian blur;
[0023] Step 103: Use a computer to record both the original power line image and the enhanced power line image as power line images.
[0024] For the above power line foreign object detection method based on image vision, further, step three is as follows:
[0025] Step 301: Use a computer to input the labeled power line image into the backbone layer of the backbone network for feature extraction to obtain the first power line feature map.
[0026] Step 302: Use a computer to input the first power line feature map into the first stage layer for feature extraction to obtain the second power line feature map.
[0027] Step 303: Use a computer to input the second power line feature map into the second stage layer for feature extraction to obtain the third power line feature map. At the same time, the third power line feature map passes through the first SKAttention module and outputs the first intermediate feature map.
[0028] Step 304: Use a computer to input the third power line feature map into the third stage layer for feature extraction to obtain the fourth power line feature map. At the same time, the fourth power line feature map passes through the second SKAttention module and outputs the second intermediate feature map.
[0029] Step 305: Use a computer to input the fourth power line feature map through the fourth stage layer, the third SKAttention module, and the SPPF module in sequence for feature extraction to obtain the fifth power line feature map.
[0030] Step 306: Use a computer to obtain the first downsampled feature map by passing the second intermediate feature map through the first downsampling module. Perform image addition processing on the first downsampled feature map and the fifth power line feature map to obtain the first fusion feature map.
[0031] Use a computer to obtain the first upsampled feature map by passing the first fusion feature map through the first upsampling module. Concatenate the first upsampled feature map and the second intermediate feature map through Concat to obtain the first concatenated feature map. Pass the first concatenated feature map through the first C2f module to obtain the first concatenated extraction feature map.
[0032] Step 307: Use a computer to obtain the second downsampled feature map by passing the first intermediate feature map through the second downsampling module. Perform image addition processing on the second downsampled feature map and the first concatenated extraction feature map to obtain the second fusion feature map.
[0033] The computer is used to obtain a second upsampled feature map by passing the second fused feature map through a second upsampling module; the second upsampled feature map and the first intermediate feature map are concatenated through Concat to obtain a second concatenated feature map; the second concatenated feature map is output as a first feature map through a second C2f module;
[0034] Step 308: The computer is used to concatenate the feature map after the first feature map passes through the first CBS module and the first concatenated extraction feature map through Concat to obtain a third concatenated feature map; the third concatenated feature map is output as a second feature map through a third C2f module;
[0035] Step 309: The computer is used to concatenate the feature map after the second feature map passes through the second CBS module and the fifth transmission line feature map through Concat to obtain a fourth concatenated feature map; the fourth concatenated feature map is output as a third feature map through a fourth C2f module.
[0036] In the above method for detecting foreign objects on transmission lines based on image vision, further, the Head module includes a first detection head, a second detection head, and a third detection head, and the first detection head, the second detection head, and the third detection head each include two CBS modules and a Conv2d module;
[0037] Step four, the specific process is as follows:
[0038] Step 401: The computer is used to input the first feature map into the first detection head for processing to obtain a first detection feature map; the computer is used to input the second feature map into the second detection head for processing to obtain a second detection feature map; the computer is used to input the third feature map into the third detection head for processing to obtain a third detection feature map;
[0039] Step 402: Select the first four channel feature maps from the first detection feature map, the second detection feature map, and the third detection feature map as the x-direction offset t x of the center point relative to the grid cell, the y-direction offset t y , the width of the prediction box, and the height of the prediction box;
[0040] Step 403: According to obtain the x coordinate X of the center point of the prediction box and the y coordinate Y of the center point of the prediction box; where S represents the feature map reduction factor, and W C represents the length and width of the transmission line image, W C1 represents the length of the first detection feature map, the second detection feature map, or the third detection feature map, σ(·) represents the Sigmoid function, L x represents the x coordinate of the grid cell on the feature map, and L y represents the y coordinate of the grid cell on the feature map;
[0041] Step 404: Obtain the confidence output value Cout of the prediction box from the fifth channel feature map, and based on obtain the confidence value of the prediction box
[0042] Step 405: Obtain the corresponding classification value P of the grid cell from the last 2 channel feature maps, and based on p = σ(P), obtain the class probability p of the prediction box.
[0043] For the above method for detecting foreign objects on transmission lines based on image vision, further, in step 501, the specific process is as follows:
[0044] Step 5011: Use a computer to obtain the number of pixels S′ in the intersection region between the prediction box and the true annotation box in the annotated transmission line image, and use the computer to calculate according to to obtain the intersection over union IoU between the prediction box and the true annotation box; where h represents the height of the true annotation box, w represents the width of the true annotation box; h′ represents the height of the prediction box, and w′ represents the width of the prediction box;
[0045] Step 5012: Denote the prediction boxes with IoU greater than 0.6 and the prediction box corresponding to the maximum IoU as positive sample prediction boxes, and denote the prediction boxes with IoU less than 0.3 as negative sample prediction boxes;
[0046] Step 5013: Use a computer to calculate according to to obtain the similarity measure v of the aspect ratio of the i'-th positive sample prediction box i′ ; where π represents pi; h i ″ represents the height of the i'-th positive sample prediction box, w i ″ represents the width of the i'-th positive sample prediction box;
[0047] Step 5014: Use a computer to calculate according to to obtain the optimized bounding box regression loss L box ; where IoU i′ represents the intersection over union between the i'-th positive sample prediction box and the true annotation box, ρ i′ represents the Euclidean distance between the center point of the i'-th positive sample prediction box and the center point of the true annotation box, c i′ represents the diagonal length of the minimum bounding box between the i'-th positive sample prediction box and the true annotation box, α is a weight coefficient, and L1(i′) represents the L1 loss between the i'-th positive sample prediction box and the true annotation box; N z is the total number of positive sample prediction boxes; i′ and N z are positive integers, and 1 ≤ i′ ≤ N z ;
[0048] Step 5015. Use a computer to calculate according to to obtain the confidence loss L conf ; where, if the i-th prediction box is a positive sample prediction box, then y i takes the value of 1, if the i-th prediction box is a negative sample prediction box, then y i takes the value of 0, represents the confidence value of the i-th prediction box; N is the total number of positive sample prediction boxes and negative sample prediction boxes, i and N are positive integers, and 1 ≤ i ≤ N;
[0049] Step 5016. Use a computer to calculate according to to obtain the classification loss L cls ; where, y i′ represents the label of the i'-th positive sample prediction box, and y i′ takes the value of 1, p i′ represents the class probability of the i'-th positive sample prediction box; N z is the total number of positive sample prediction boxes.
[0050] The present invention has the following advantages compared with the prior art:
[0051] 1. The method steps of the present invention are simple and reasonably designed, which solves the problem of detecting foreign objects on transmission lines, namely bird nests and kites.
[0052] 2. The present invention sets up the SKAttention module, and the output of the C2f module in the second stage layer is connected to the first SKAttention module, the output of the C2f module in the third stage layer is connected to the second SKAttention module, and the output of the C2f module in the fourth stage layer is connected to the third SKAttention module. By connecting the output of each C2f module to the SKAttention module, the representation ability of the model is enhanced in the feature extraction stage, and it can help the model better focus on the key areas in the image, thereby improving the detection performance.
[0053] 3. The present invention also introduces a bidirectional feature fusion module in the YOLOv8 model, performs downsampling addition operation on the feature map output by the backbone network, and then performs upsampling after addition. Through bidirectional feature fusion of downsampling and upsampling, the receptive field can be increased, and the details can be better captured. Through multi-level feature fusion, the information loss or degradation in the process of feature transmission and interaction can be prevented, the scale of the target can be better matched, and the detection accuracy and effect can be improved.
[0054] 4. The present invention distinguishes positive sample prediction boxes and negative sample prediction boxes through the intersection over union, so that only the positive sample prediction boxes are used for the boundary regression loss and classification loss in the subsequent process, which can reduce the calculation amount, speed up the training speed, and can focus more on the features of the target, thereby improving the detection accuracy.
[0055] 5. In the optimization of the boundary regression loss of the present invention, not only the GIoU loss is considered, but also the L1 loss between the predicted bounding box of the positive sample and the ground truth bounding box is added, which can accelerate the model convergence, improve the accuracy of the bounding box, and enhance the robustness of the model.
[0056] In summary, the method steps of the present invention are simple and reasonably designed. The foreign objects in the transmission line image are labeled by the minimum circumscribed rectangle, and based on the YOLOv8 network model, the SKAttention module and the bidirectional feature fusion module are added to extract and detect the features of the foreign objects in the transmission line image. The bidirectional feature fusion module fuses the multi-level features, and the SKAttention module strengthens the key features to improve the detection accuracy.
[0057] Next, through the attached drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings
[0058] Figure 1 It is a flow chart of the method of the present invention.
[0059] Figure 2 It is a model block diagram of the present invention.
[0060] Figure 3 It is a detection diagram of the foreign object kite of the present invention.
[0061] Figure 4 It is a detection diagram of the foreign object bird's nest of the present invention. Detailed Embodiments
[0062] As Figure 1 and Figure 2 shown, the method for detecting foreign objects on transmission lines based on image vision of the present invention includes the following steps:
[0063] Step 1. Acquisition and annotation of transmission line images:
[0064] Acquire the transmission line image, and use the computer to label the foreign objects in the transmission line image with the minimum circumscribed rectangle using the labelimg annotation software to obtain the annotated transmission line image, which is recorded as the training set; wherein, a minimum circumscribed rectangle in the annotated transmission line image represents a ground truth bounding box, and the marking information of each ground truth bounding box includes the x coordinate, y coordinate, box width, and box height of the center point; wherein, the foreign objects are bird's nests and kites.
[0065] Step 2. Construct an improved YOLOv8 network model for detecting foreign objects on transmission lines:
[0066] The improved YOLOv8 network model for detecting foreign objects on transmission lines includes a YOLOv8 model and an SKAttention module; the YOLOv8 model includes a backbone network, a bidirectional feature fusion module, and a Head module. The backbone network includes a backbone layer, a first-stage layer, a second-stage layer, a third-stage layer, a fourth-stage layer, and an SPPF module. The backbone layer is a CBS module, and the first-stage layer to the fourth-stage layer each include a CBS module and a C2f module; the SKAttention module includes a first SKAttention module, a second SKAttention module, and a third SKAttention module. The output of the C2f module in the second-stage layer is connected to the first SKAttention module, the output of the C2f module in the third-stage layer is connected to the second SKAttention module, and a third SKAttention module is connected between the C2f module in the fourth-stage layer and the SPPF module;
[0067] The bidirectional feature fusion module includes a first upsampling module, a second upsampling module, a first downsampling module, a second downsampling module, a first C2f module, a second C2f module, a third C2f module, a fourth C2f module, a first CBS module, and a second CBS module;
[0068] Step 3: Feature extraction of the transmission line image:
[0069] The labeled transmission line image is subjected to feature extraction through the improved YOLOv8 network model for detecting foreign objects on transmission lines to obtain a first feature map, a second feature map, and a third feature map;
[0070] Step 4: Prediction after feature extraction of the transmission line image:
[0071] The computer processes and detects the first feature map, the second feature map, and the third feature map respectively through the Head module. The x coordinate X of the center point of the prediction box, the y coordinate Y of the center point of the prediction box, the width w' of the prediction box, and the height h' of the prediction box are obtained through the first four channel feature maps; the confidence value of the prediction box is obtained through the fifth channel feature map; the classification probability of the prediction box is obtained through the last two channel feature maps;
[0072] Step 5: Training of the improved YOLOv8 network model for detecting foreign objects on transmission lines:
[0073] Step 501: Input the x coordinate X of the center point of the prediction box, the y coordinate Y of the center point of the prediction box, the width w' of the prediction box, the height h' of the prediction box, the confidence value, and the classification probability in Step 4. The computer calculates according to L loss = 0.5L box + L conf + 0.1L cls, the total loss function L is obtained loss ; where L box represents the optimized boundary regression loss, and L conf represents the confidence loss, and L cls represents the classification loss;
[0074] Step 502: The computer uses the SGD optimizer and utilizes the total loss function L loss to iteratively train the improved YOLOv8 network model for overhead transmission line foreign object detection until all the training sets are trained, completing one iteration of training;
[0075] Step 503: Repeat Step 502 for iterative training until the preset number of iterations is satisfied, obtaining the trained improved YOLOv8 network model for overhead transmission line foreign object detection;
[0076] Step Six: Use the trained improved YOLOv8 network model for overhead transmission line foreign object detection to detect foreign objects in subsequent overhead transmission line images.
[0077] In this embodiment, the acquisition of the overhead transmission line image in Step One is specifically as follows:
[0078] Step 101: Use a drone equipped with a camera to collect images of the overhead transmission line to be detected, obtain the overhead transmission line image, and transmit it to the computer to obtain the original overhead transmission line image;
[0079] Step 102: Use the computer to perform image processing on the original overhead transmission line image to obtain the enhanced overhead transmission line image; where the image processing includes vertical flipping, horizontal flipping, random brightness, affine transformation, contrast enhancement, or Gaussian blur;
[0080] Step 103: Use the computer to record both the original overhead transmission line image and the enhanced overhead transmission line image as the overhead transmission line image.
[0081] In this embodiment, Step Three is specifically as follows:
[0082] Step 301: Use the computer to input the labeled overhead transmission line image into the backbone layer of the backbone network for feature extraction to obtain the first overhead transmission line feature map;
[0083] Step 302: Use the computer to input the first overhead transmission line feature map into the first stage layer for feature extraction to obtain the second overhead transmission line feature map;
[0084] Step 303: Use the computer to input the second overhead transmission line feature map into the second stage layer for feature extraction to obtain the third overhead transmission line feature map; meanwhile, the third overhead transmission line feature map passes through the first SKAttention module and outputs the first intermediate feature map;
[0085] Step 304: Use a computer to input the third transmission line feature map into the third stage layer for feature extraction to obtain the fourth transmission line feature map; meanwhile, the fourth transmission line feature map passes through the second SKAttention module and outputs the second intermediate feature map;
[0086] Step 305: Use a computer to successively pass the fourth transmission line feature map through the fourth stage layer, the third SKAttention module, and the SPPF module for feature extraction to obtain the fifth transmission line feature map;
[0087] Step 306: Use a computer to obtain the first downsampled feature map by passing the second intermediate feature map through the first downsampling module; perform image addition processing on the first downsampled feature map and the fifth transmission line feature map to obtain the first fused feature map; Use a computer to obtain the first upsampled feature map by passing the first fused feature map through the first upsampling module; concatenate the first upsampled feature map and the second intermediate feature map through Concat to obtain the first concatenated feature map; pass the first concatenated feature map through the first C2f module to obtain the first concatenated extraction feature map;
[0088] Use a computer to obtain the second downsampled feature map by passing the first intermediate feature map through the second downsampling module; perform image addition processing on the second downsampled feature map and the first concatenated extraction feature map to obtain the second fused feature map; Use a computer to obtain the second upsampled feature map by passing the second fused feature map through the second upsampling module; concatenate the second upsampled feature map and the first intermediate feature map through Concat to obtain the second concatenated feature map; output the first feature map by passing the second concatenated feature map through the second C2f module;
[0089] Step 307: Use a computer to obtain the second downsampled feature map by passing the first intermediate feature map through the second downsampling module; perform image addition processing on the second downsampled feature map and the first concatenated extraction feature map to obtain the second fused feature map; Use a computer to obtain the second upsampled feature map by passing the second fused feature map through the second upsampling module; concatenate the second upsampled feature map and the first intermediate feature map through Concat to obtain the second concatenated feature map; output the first feature map by passing the second concatenated feature map through the second C2f module;
[0090] Use a computer to obtain the second upsampled feature map by passing the second fused feature map through the second upsampling module; concatenate the second upsampled feature map and the first intermediate feature map through Concat to obtain the second concatenated feature map; output the first feature map by passing the second concatenated feature map through the second C2f module;
[0091] Step 308: Use a computer to concatenate the feature map after passing the first feature map through the first CBS module and the first concatenated extraction feature map through Concat to obtain the third concatenated feature map; output the second feature map by passing the third concatenated feature map through the third C2f module;
[0092] Step 309: Use a computer to concatenate the feature map after passing the second feature map through the second CBS module and the fifth transmission line feature map through Concat to obtain the fourth concatenated feature map; output the third feature map by passing the fourth concatenated feature map through the fourth C2f module. In this embodiment, the Head module includes a first detection head, a second detection head, and a third detection head, and the first detection head, the second detection head, and the third detection head each include two CBS modules and a Conv2d module;
[0093] In this embodiment, the Head module includes a first detection head, a second detection head, and a third detection head, and the first detection head, the second detection head, and the third detection head each include two CBS modules and a Conv2d module;
[0094] Step four, the specific process is as follows:
[0095] Step 401: Use a computer to input the first feature map into the first detection head for processing to obtain a first detection feature map; use a computer to input the second feature map into the second detection head for processing to obtain a second detection feature map; use a computer to input the third feature map into the third detection head for processing to obtain a third detection feature map;
[0096] Step 402: Select the first four channel feature maps from the first detection feature map, the second detection feature map, and the third detection feature map as the x-direction offset t of the center point relative to the grid cell x , the y-direction offset t y , the width of the prediction box, and the height of the prediction box of the prediction box;
[0097] Step 403: According to obtain the x coordinate X of the center point of the prediction box and the y coordinate Y of the center point of the prediction box; where S represents the feature map reduction factor, and W C represents the length and width of the transmission line image, W C1 represents the length of the first detection feature map, the second detection feature map, or the third detection feature map, σ(·) represents the Sigmoid function, L x represents the x coordinate of the grid cell on the feature map, L y represents the y coordinate of the grid cell on the feature map;
[0098] Step 404: Obtain the confidence output value Cout of the prediction box through the fifth channel feature map, and according to obtain the confidence value of the prediction box
[0099] Step 405: Obtain the corresponding classification value P of the grid cell through the last 2 channel feature maps, and according to p = σ(P), obtain the class probability p of the prediction box.
[0100] In this embodiment, step 501 is specifically as follows:
[0101] Step 5011: Use a computer to obtain the number of pixels S′ in the intersection area between the prediction box and the true annotation box in the annotated transmission line image, and use a computer to according to obtain the intersection over union IoU between the prediction box and the true annotation box; where h represents the height of the true annotation box, w represents the width of the true annotation box; h′ represents the height of the prediction box, and w′ represents the width of the prediction box;
[0102] Step 5012: Denote the prediction box with an IoU greater than 0.6 and the prediction box corresponding to the maximum IoU as positive sample prediction boxes, and denote the prediction box with an IoU less than 0.3 as a negative sample prediction box;
[0103] Step 5013: Use a computer to calculate according to to obtain the similarity measure v of the aspect ratio of the i'-th positive sample prediction box i′ ; where, π represents pi; h i ″ represents the height of the i'-th positive sample prediction box, and w i ″ represents the width of the i'-th positive sample prediction box;
[0104] Step 5014: Use a computer to calculate according to to obtain the optimized bounding box regression loss L box ; where, IoU i′ represents the intersection over union of the i'-th positive sample prediction box and the ground truth box, ρ i′ represents the Euclidean distance between the center point of the i'-th positive sample prediction box and the center point of the ground truth box, c i′ represents the diagonal length of the minimum bounding box between the i'-th positive sample prediction box and the ground truth box, α is a weight coefficient, and L1(i′) represents the L1 loss between the i'-th positive sample prediction box and the ground truth box; N z is the total number of positive sample prediction boxes; i′ and N z are positive integers, and 1 ≤ i′ ≤ N z ;
[0105] Step 5015: Use a computer to calculate according to to obtain the confidence loss L conf ; where, if the i-th prediction box is a positive sample prediction box, then y i takes the value of 1, if the i-th prediction box is a negative sample prediction box, then y i takes the value of 0, represents the confidence value of the i-th prediction box; N is the total number of positive sample prediction boxes and negative sample prediction boxes, i and N are positive integers, and 1 ≤ i ≤ N;
[0106] Step 5016: Use a computer to calculate according to to obtain the classification loss L cls ; where, y i′ represents the label of the i'-th positive sample prediction box, and y i′ takes the value of 1, p i′ represents the class probability of the i'-th positive sample prediction box; N z is the total number of positive sample prediction boxes.
[0107] In this embodiment, the SKAttention module, namely the Selective Kernel Attention module, is the SK attention mechanism module; the SPPF module, namely the Spatial Pyramid Pooling Fast module, is the fast spatial pyramid pooling module; the C2f module is a hybrid neural network structure that combines a convolutional neural network and a fully connected neural network to learn global features while retaining local features of the image.
[0108] In this embodiment, the backbone layer is the Conv (convolutional layer). The CBS module is mainly composed of three parts: Conv (convolutional layer), BN (Batch Normalization), and SiLU (activation function). The size of the Conv (convolutional layer) is 3×3, the stride is 2, and the padding is 1.
[0109] In this embodiment, in actual use, the size of the power line image is represented by length × width × number of channels. After annotation, the size of the power line image is 640×640×3; the number of convolutional kernels in the backbone layer is 64, and the size of the first power line feature map is 320×320×64; the number of convolutional kernels in the first stage layer is 128, and the size of the second power line feature map is 160×160×128; the number of convolutional kernels in the second stage layer is 256, and the size of the third power line feature map is 80×80×256; the number of convolutional kernels in the third stage layer is 512, and the size of the fourth power line feature map is 40×40×512; the number of convolutional kernels in the fourth stage layer is 512, and the size of the fifth power line feature map is 20×20×512.
[0110] The size of the second intermediate feature map is 40×40×512, the size of the first downsampled feature map is 20×20×512, the size of the first fused feature map is 20×20×512, the size of the first upsampled feature map is 40×40×512, the size of the first concatenated feature map is 40×40×1024, and the size of the first concatenated and extracted feature map is 40×40×256;
[0111] The size of the first intermediate feature map is 80×80×256, the size of the second downsampled feature map is 40×40×256, the size of the second fused feature map is 40×40×256, the size of the second upsampled feature map is 80×80×256, the size of the second concatenated feature map is 80×80×512, and the size of the first feature map is 80×80×256;
[0112] The size of the third concatenated feature map is 40×40×512, and the size of the second feature map is 40×40×512;
[0113] The size of the fourth spliced feature map is 20×20×1024, and the size of the third feature map is 20×20×512.
[0114] In this embodiment, during actual use, the size of the Conv (convolutional layer) of the CBS module in the first detection head, the second detection head, and the third detection head is 3×3, the stride is 1, the padding is 1, and the number of convolutional kernels is 128; the number of convolutional kernels in the Conv2d module is 7, the size of the convolutional kernel is 1×1, the stride is 1, and the padding is 0. Then the size of the first detection feature map is 80×80×7, the size of the second detection feature map is 40×40×7, and the size of the third detection feature map is 20×20×7.
[0115] In this embodiment, during actual use, the length and width of the image are the same.
[0116] In this embodiment, during actual use, the trained improved YOLOv8 network model for overhead transmission line foreign object detection is also tested with a test set to meet the requirements.
[0117] In this embodiment, during actual use, the x coordinate X of the center point of the prediction box and the y coordinate Y of the center point of the prediction box are rounded.
[0118] In this embodiment, during actual use, the x-direction offset t x of the center point relative to the grid cell and the y-direction offset t y are processed by the Sigmoid function to be restricted within the range of [0,1] to ensure that the center point of the bounding box is within the current grid cell. By restricting the offset range, the model stability is improved.
[0119] In this embodiment, during actual use, the preset number of iterative training is 50 - 100.
[0120] In this embodiment, during actual use, such as Figure 3 and Figure 4 , it can achieve the detection of foreign objects such as bird nests and kites on the overhead transmission line image and meet the requirements of overhead transmission line foreign object detection.
[0121] In summary, the method of the present invention has simple steps and reasonable design. The foreign objects in the overhead transmission line image are marked by the minimum bounding rectangle, and the SKAttention module and the bidirectional feature fusion module are added based on the YOLOv8 network model to extract and detect the features of the foreign objects in the overhead transmission line image. The bidirectional feature fusion module fuses the multi-level features, and the SKAttention module strengthens the key features to improve the detection accuracy.
[0122] The above are only the preferred embodiments of the present invention, and do not impose any limitations on the present invention. Any simple modifications, changes, and equivalent structural changes made to the above embodiments based on the technical essence of the present invention still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for detecting foreign matter in a transmission line based on image vision, characterized in that: The method comprises the following steps: Step 1: Acquisition and annotation of transmission line images: Acquire a transmission line image, and use a computer to use labelimg labeling software to label foreign objects in the transmission line image with a minimum circumscribed rectangular frame to obtain a labeled transmission line image, and record it as a training set; wherein a minimum circumscribed rectangular frame in the labeled transmission line image represents a true labeling frame, and the labeling information of each true labeling frame includes the x-coordinate of the center point, the y-coordinate, the frame width and the frame height; wherein the foreign objects are bird nests and kites; Step 2: Build an improved YOLOv8 network model for foreign body detection on power transmission lines: The improved YOLOv8 network model for foreign body detection on power transmission lines includes a YOLOv8 model and a SKAttention module; the YOLOv8 model includes a backbone network, a bidirectional feature fusion module and a Head module, the backbone network includes a backbone layer, a first stage layer, a second stage layer, a third stage layer, a fourth stage layer and an SPPF module, the backbone layer is a CBS module, and the first stage layer to the fourth stage layer all include a CBS module and a C2f module; the SKAttention module includes a first SKAttention module, a second SKAttention module and a third SKAttention module, the output of the C2f module of the second stage layer is connected to the first SKAttention module, the output of the C2f module of the third stage layer is connected to the second SKAttention module, and the third SKAttention module is connected between the C2f module of the fourth stage layer and the SPPF module; The bidirectional feature fusion module includes a first upsampling module, a second upsampling module, a first downsampling module, a second downsampling module, a first C2f module, a second C2f module, a third C2f module, a fourth C2f module, a first CBS module and a second CBS module; Step 3: Feature extraction of transmission line images: The annotated transmission line image is subjected to feature extraction through the transmission line foreign body detection improved YOLOv8 network model to obtain the first feature map, the second feature map and the third feature map; Step 4: Prediction after feature extraction of transmission line image: The first feature map, the second feature map and the third feature map are processed and detected by the Head module by a computer, and the center point x coordinate X of the prediction box, the center point y coordinate Y of the prediction box, the width w′ of the prediction box and the height h′ of the prediction box are obtained through the first four channel feature maps; the confidence value of the prediction box is obtained through the fifth channel feature map; and the classification probability of the prediction box is obtained through the last two channel feature maps; Step 5: Training of improved YOLOv8 network model for foreign body detection on power transmission lines: Step 501: Input the x-coordinate X of the center point of the prediction box, the y-coordinate Y of the center point of the prediction box, the width w′ of the prediction box, the height h′ of the prediction box, the confidence value and the classification probability in step 4. loss =0.5L box +L conf +0.1L cls , and get the total loss function L loss Among them, L box represents the optimization boundary regression loss, L conf represents the confidence loss, L cls represents the classification loss; Step 502: The computer uses the SGD optimizer to use the total loss function L loss Iterative training is performed on the improved YOLOv8 network model for foreign body detection on power transmission lines until all training sets are trained and one iteration of training is completed; Step 503, repeating the iterative training in step 502 until the preset number of iterative training times is met, and obtaining a trained improved YOLOv8 network model for foreign body detection on power transmission lines; Step 6: Use the trained transmission line foreign object detection to improve the YOLOv8 network model to perform foreign object detection on subsequent transmission line images.
2. A method for detecting foreign matter in a power transmission line based on image vision according to claim 1, characterized in that: The specific process of acquiring the transmission line image in step 1 is as follows: Step 101: Use a camera mounted on a drone to collect images of the transmission line to be inspected, obtain a transmission line image, and transmit the image to a computer to obtain an original transmission line image; Step 102: using a computer to process the original transmission line image to obtain an enhanced transmission line image; wherein the image processing includes vertical flipping, horizontal flipping, random brightness, affine transformation, contrast enhancement or Gaussian blur; Step 103: Using a computer, record both the original transmission line image and the enhanced transmission line image as transmission line images.
3. A method for detecting foreign matter in a power transmission line based on image vision according to claim 1, characterized in that: Step 3: The specific process is as follows: Step 301: Use a computer to input the annotated transmission line image into the backbone layer of the backbone network for feature extraction to obtain a first transmission line feature map; Step 302: Using a computer, input the first transmission line feature map into the first stage layer for feature extraction to obtain a second transmission line feature map; Step 303: Use a computer to input the second transmission line feature map into the second stage layer for feature extraction to obtain a third transmission line feature map; at the same time, the third transmission line feature map passes through the first SKAttention module to output a first intermediate feature map; Step 304: Use a computer to input the third transmission line feature map into the third stage layer for feature extraction to obtain a fourth transmission line feature map; at the same time, the fourth transmission line feature map passes through the second SKAttention module to output a second intermediate feature map; Step 305: Use a computer to extract features from the fourth transmission line feature map through the fourth stage layer, the third SKAttention module, and the SPPF module in sequence to obtain a fifth transmission line feature map; Step 306: Use a computer to process the second intermediate feature map through a first downsampling module to obtain a first downsampling feature map; perform image addition processing on the first downsampling feature map and the fifth power transmission line feature map to obtain a first fusion feature map; A computer is used to pass the first fusion feature map through a first upsampling module to obtain a first upsampling feature map; the first upsampling feature map is concatenated with the second intermediate feature map through Concat to obtain a first concatenated feature map; the first concatenated feature map is passed through a first C2f module to obtain a first concatenated extracted feature map; Step 307: using a computer to process the first intermediate feature map through a second downsampling module to obtain a second downsampling feature map; performing image addition processing on the second downsampling feature map and the first spliced extracted feature map to obtain a second fused feature map; A computer is used to pass the second fused feature map through a second upsampling module to obtain a second upsampling feature map; the second upsampling feature map is concatenated with the first intermediate feature map through Concat to obtain a second concatenated feature map; the second concatenated feature map is passed through a second C2f module to output the first feature map; Step 308: Use a computer to concatenate the feature map after the first feature map passes through the first CBS module with the first splicing extraction feature map through Concat to obtain a third splicing feature map; and output the third splicing feature map through a third C2f module as a second feature map; Step 309, concatenate the feature map of the second feature map of the computer after passing through the second CBS module with the fifth power transmission line feature map through Concat to obtain a fourth concatenated feature map; and output the fourth concatenated feature map through the fourth C2f module to the third feature map.
4. A method for detecting foreign matter in a power transmission line based on image vision according to claim 1, characterized in that: The Head module includes a first detection head, a second detection head and a third detection head, and the first detection head, the second detection head and the third detection head each include two CBS modules and a Conv2d module; Step 4: The specific process is as follows: Step 401: Use a computer to input the first feature map into the first detection head for processing to obtain a first detection feature map; use a computer to input the second feature map into the second detection head for processing to obtain a second detection feature map; use a computer to input the third feature map into the third detection head for processing to obtain a third detection feature map; Step 402: Select the first four channel feature maps from the first detection feature map, the second detection feature map, and the third detection feature map as the x-direction offset t of the center point relative to the grid unit. x , y-direction offset t y , the width of the prediction box and the height of the prediction box; Step 403: According to Get the x-coordinate X of the center point of the prediction box and the y-coordinate Y of the center point of the prediction box; where S represents the feature map reduction coefficient, and W C represents the length and width of the transmission line image, W C1 represents the length of the first detection feature map, the second detection feature map, or the third detection feature map, σ(·) represents the Sigmoid function, and L x represents the x-coordinate of the grid cell on the feature map, L y Represents the y coordinate of the grid cell on the feature map; Step 404: Obtain the confidence output value Cout of the prediction box through the fifth channel feature map, and Get the confidence value of the predicted box Step 405: Obtain the corresponding classification value P of the grid unit through the latter 2-channel feature map, and obtain the category probability p of the prediction box according to p=σ(P).
5. A method for detecting foreign matter in a power transmission line based on image vision according to claim 4, characterized in that: Step 501, the specific process is as follows: Step 5011: Use a computer to obtain the number of pixels S′ in the intersection area between the predicted box and the real marked box in the annotated transmission line image. Use a computer according to Get the intersection over union (IoU) of the predicted box and the true labeled box; where h represents the height of the true labeled box, w represents the width of the true labeled box; h′ represents the height of the predicted box, and w′ represents the width of the predicted box; Step 5012: The prediction box with an IoU greater than 0.6 and the prediction box corresponding to the maximum IoU value are recorded as positive sample prediction boxes, and the prediction box with an IoU less than 0.3 is recorded as a negative sample prediction box; Step 5013: Using a computer according to Get the similarity measure v of the aspect ratio of the i′th positive prediction box i′ ; Among them, π represents pi; h' i′ represents the height of the i′th positive prediction box, w′ i′ Represents the width of the i′th positive prediction box; Step 5014: Using a computer according to Get the optimized boundary regression loss L box Among them, IoU i′ represents the intersection-over-union ratio of the i′th positive prediction box and the true annotation box, ρ i′ represents the Euclidean distance between the center point of the predicted box of the i′th positive sample and the center point of the true annotation box, c i′ represents the diagonal length of the minimum bounding box between the i′th positive sample prediction box and the true annotation box, α is the weight coefficient, and L1(i′) represents the L1 loss between the predicted box of the i′th positive sample and the true labeled box; N z is the total number of positive sample prediction boxes; i′ and N z is a positive integer, and 1≤i′≤N z ; Step 5015: Using a computer according to Get the confidence loss L conf ; If the i-th prediction box is a positive sample prediction box, then y i The value is 1. If the i-th prediction box is a negative sample prediction box, then y i The value is 0. Represents the confidence value of the i-th prediction box; N is the total number of positive sample prediction boxes and negative sample prediction boxes, i and N are positive integers, and 1≤i≤N; Step 5016: Using a computer according to Get the classification loss L cls ; Among them, y i′ represents the label of the i′th positive sample prediction box, and y i′ The value is 1, p i′ Represents the category probability of the i′th positive sample prediction box; N z is the total number of positive sample prediction boxes.
Citation Information
Cited By
Power transmission line inspection image environment adaptive target detection method and system
CN122157066A