Power transmission line foreign matter invasion defect detection method based on improved YOLOv7 network

By improving the YOLOv7 network and combining transfer learning and data enhancement technology, the problems of low efficiency and poor recognition effect of bird nest detection on the transmission line tower are solved, and high-accuracy and efficient bird nest detection effect are achieved.

CN119992381APending Publication Date: 2025-05-13INFORMATION & TELECOMM COMPANY SICHUAN ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510065058.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is inefficient and difficult in the detection of bird nests of power transmission line pole towers, and the recognition effect is not good during drone inspection.

Method used

Using the improved YOLOv7 network, the model is trained to improve the accuracy of bird nest recognition through transfer learning and data augmentation methods. The specific steps include building an improved YOLOv7 network, creating virtual data sets using Unreal 5, performing transfer learning training, and deploying the model to the drone recognition device.

Benefits of technology

It realizes high accuracy and efficient detection of bird nests on the power transmission line poles and towers, solving the problems of low traditional manual detection efficiency and poor drone recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992381A_ABST
    Figure CN119992381A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image recognition, and discloses an improved YOLOv7 network-based power transmission line foreign matter intrusion defect detection method, which comprises the following steps of: carrying out transfer learning by utilizing a power transmission line bird nest data set to train the improved YOLOv7 network, deploying the trained YOLOv7 network on recognition equipment, uploading a shot image to the recognition equipment, and identifying the foreign matter intrusion defect of the power transmission line. And the bird nest is detected. Intelligent recognition of the bird nest of the power transmission line tower is achieved, and the problems that manual detection of the bird nest is low in efficiency, large in difficulty and the like are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image recognition technology, and relates to the detection of foreign body intrusion into power transmission lines, and in particular to the detection of bird nests on power transmission line towers. The present invention uses a camera carried by an unmanned aerial vehicle to collect images of towers at various angles, and then uses an improved YOLOV7 network to detect bird nests in the images. Background Art

[0002] Transmission line foreign bodies (such as bird nests) may cause multiple safety hazards when they are formed on transmission line towers. First, bird nests as conductors may cause power line tripping. Especially on rainy days, when the materials in the bird nests come into contact with the transmission lines, the current may pass through the bird nests instead of the normal conductor path, causing the line to trip and interrupt the power supply. This will not only cause inconvenience to the power supply system, but may also affect the normal power consumption of users. Secondly, the branches of the bird nests are prone to catch fire or fall on the insulators in a dry environment, causing insulator flashover. The materials in the bird nests may be flammable under high temperature and dry conditions. Once a fire occurs, it may cause a failure of the transmission line. At the same time, the branches in the bird nests may fall on the insulators, destroy the insulation performance, cause insulator flashover, and then affect the normal operation of the line. Insulator flashover not only causes energy loss, but also may cause equipment damage or arc discharge, posing a threat to the safety of personnel and equipment. In addition, the frequent activities of birds near the nests increase the probability of high-voltage transmission line failures. Birds may fly, roost or nest near bird nests, and their activities may cause collisions, short circuits or other unexpected situations, which may lead to line faults. These faults not only cause power outages, but may also cause damage to transmission line equipment, increasing the cost of repair and restoration of power.

[0003] Therefore, bird nest detection on transmission line towers is an important task based on the concern for the safety of transmission lines. It is of great significance to ensure the safety of transmission lines and the stable operation of power grids.

[0004] At present, the traditional manual inspection of bird nests on power transmission lines has problems such as low efficiency and difficulty, and it takes a lot of manpower and material resources to eliminate bird nests. In recent years, with the development of drone technology, the detection of bird nests by drones has gradually become a major detection method due to its high efficiency, low risk and low cost. During drone inspections, it is necessary to carry detection equipment to detect bird nests using target detection algorithms. Although it can effectively replace human inspections, there are still some problems. First, since the proportion of bird nests in the transmission line tower images taken by drones is small and they are severely blocked by the tower body, the recognition accuracy is low; secondly, due to the small number of transmission line tower images with bird nests, the training of the target detection model is not sufficient, and the detection effect of the model is not good.

[0005] In summary, in order to ensure the safe operation of transmission lines, timely detection and treatment of bird nests can reduce safety risks and ensure the stability of power supply. It is very important to detect bird nests on transmission line towers. Summary of the invention

[0006] The purpose of the present invention is to provide a transmission line foreign body intrusion defect detection method based on an improved YOLOv7 network to address the problems of low efficiency and difficulty in manual detection of bird nests on transmission line towers, and poor recognition effect during drone inspections, which can achieve high-accuracy and efficient detection of bird nests on transmission line towers.

[0007] The inventive idea of ​​the present invention is: first, Unreal 5 can be used to create a virtual data set and a data enhancement method to expand the data set, and a transfer learning training model can be used; then, advanced technologies such as drones and target detection can be used for inspection to improve detection efficiency; then, in view of the small proportion of bird nests invaded by foreign objects in power transmission lines, severe occlusion, and low recognition accuracy, the YOLOv7 network is improved to improve the recognition accuracy; finally, the inspection results are uploaded.

[0008] To achieve the above objectives, the present invention adopts the following technical solutions.

[0009] The present invention provides a method for detecting foreign body intrusion defects in power transmission lines based on an improved YOLOv7 network, which mainly includes the following steps:

[0010] S1 uses the power transmission line bird nest dataset for transfer learning to train the improved YOLOv7 network; step S1 includes the following sub-steps:

[0011] S11 builds an improved YOLOv7 network; uses the kmeans clustering algorithm to adjust the proportion of anchor boxes in the YOLOv7 network;

[0012] S12 used Unreal 5 software to create a virtual data set, labeled the virtual data set and the real data set, and saved the labeled labels in yolo format;

[0013] S13 first uses a virtual data set to train the YOLOv7 network. After the training is completed, transfer learning is used to continue training the real data set to complete the training of the final model.

[0014] S2 completes the deployment of the trained YOLOv7 network and the installation of drones, recognition equipment, cameras, and batteries;

[0015] S3 completes the shooting of on-site tower images to detect bird nests;

[0016] S4 uploads the processed results to the cloud and displays the results.

[0017] In the above step S1, in order to solve the problems of small proportion of bird nests and severe occlusion, an improved YOLOv7 network is constructed to improve the accuracy of the model in identifying bird nests. First, the improved YOLOv7 network is constructed, and then the constructed improved YOLOv7 network is trained using the virtual data set and the real data set respectively.

[0018] In one implementable method, based on the width and height distribution of the bird nest target box (i.e., the bounding box of the bird nest in the image) in the collected real data set containing the bird nest, the kmeans clustering algorithm is used to determine the most appropriate anchor box ratio in the YOLOv7 network, so as to optimize the YOLOv7 network's detection ability for the bird nest target.

[0019] In one implementable manner, the feature map output by the first feature extraction unit in the backbone network of the YOLOv7 network is used in subsequent operations. In a specific implementation, the YOLOv7 network includes a backbone network for feature extraction and a head network for target recognition; the backbone network is an inverted pyramid structure, including an input layer, a first feature extraction unit, a second feature extraction unit, a third feature extraction unit and a fourth feature extraction unit arranged in sequence; the first feature extraction unit, the second feature extraction unit and the third feature extraction unit are output to the head network; the features extracted by the fourth feature extraction unit are processed by the pooling unit and input to the head network; the head network processes the input features through path aggregation to obtain several layers of output. The input layer includes three convolution modules arranged in sequence; the first feature extraction unit, the second feature extraction unit, the third feature extraction unit and the fourth feature extraction unit have the same structure, including a convolution module and a multi-branch stacking module arranged in sequence. The pooling unit uses the SPPCSPC module. The convolution module includes a two-dimensional convolution layer, a BN layer and a SiLU activation function.

[0020] Furthermore, a clockwise four-scale fusion module is provided before the input layer, and the clockwise four-scale fusion module includes four different-scale two-dimensional convolution layers, a maximum pooling layer, and a splicing unit; first, the input image is convolved with four different-scale two-dimensional convolution layers, and then the maximum pooling layer is used to reduce the size to half of the original size; then, the four convolution outputs are spliced ​​in a clockwise direction to form a new image as the input of the input layer. By using convolution kernels of different sizes to extract features, the problem of bird nests being too large or too small due to changes in the shooting position of the drone can be effectively addressed.

[0021] In one possible implementation, in the head network, a CBAM attention mechanism is added after the original Conv layer. The CBAM attention mechanism is a hybrid attention mechanism that combines spatial attention and channel attention mechanisms, which can improve the ability of the YOLOv7 network to selectively focus on and process information according to different parts of the input, thereby improving the recognition effect. In a specific implementation, the head network includes a first fusion layer, a second fusion layer, a third fusion layer, a first output layer, a second output layer, and a third output layer;

[0022] In the first fusion layer, the output features of the pooling unit are processed by the convolution layer (Conv layer), the CBAM module and the upsampling module to obtain the first feature, the output features of the third feature extraction unit are processed by the convolution layer (Conv layer) and the CBAM module to obtain the second feature, and the first feature and the second feature are concatenated and then processed by the multi-branch stacking module to obtain the first fusion feature;

[0023] In the second fusion layer, the first fusion feature is processed by a two-dimensional convolution layer and an upsampling module to obtain a third feature, the output feature of the second feature extraction unit is processed by a convolution layer (Conv layer) and a CBAM module to obtain a fourth feature, and the third feature and the fourth feature are concatenated and then processed by a multi-branch stacking module to obtain a second fusion feature;

[0024] In the third fusion layer, the second fusion feature is processed by a two-dimensional convolution layer and an upsampling module to obtain a fifth feature, the output feature of the first feature extraction unit is processed by a convolution layer (Conv layer) and a CBAM module to obtain a sixth feature, and the fifth feature and the sixth feature are concatenated and then processed by a multi-branch stacking module to obtain a third fusion feature;

[0025] In the first output layer, the output feature of the third fusion feature after the transition module is concatenated with the second fusion feature and then processed by the multi-branch stacking module and the re-parameterized convolution (Repconv) layer to obtain the first output feature;

[0026] In the second output layer, the first output feature is concatenated with the output feature of the transition module and the first fusion feature, and then processed by a multi-branch stacking module and a reparameterized convolution (Repconv) layer to obtain the second output feature;

[0027] In the third output layer, the second output feature is concatenated with the output feature of the transition module and the output feature of the pooling unit, and then processed by a multi-branch stacking module and a reparameterized convolution (Repconv) layer to obtain the third output feature.

[0028] Further, the one with better recognition effect among the first output feature, the second output feature and the third output feature is taken as the final recognition result. For example, the final recognition result is selected by conventional non-maximum suppression (NMS).

[0029] In one possible implementation, the two-dimensional convolution layer in any of the above improved YOLOv7 networks is replaced by a dilated convolution layer, and the dilation factor of the dilated convolution layer is 2. The dilated convolution introduces spacing between the convolution kernels, thereby expanding the receptive field and capturing a wider range of contextual information without increasing the amount of computation.

[0030] In one possible implementation, the YOLOv7 network is improved to change the IOU (intersection over union) operation in the loss function to CIOU.

[0031] The calculation formula of CIOU is as follows:

[0032]

[0033] Among them, IOU is the intersection area of ​​the predicted box and the real box divided by the area of ​​the two together; ρ(b,b gt ) represents the Euclidean distance between the center points of the predicted box and the true box, and c represents the diagonal distance of the minimum closure area that can contain both the predicted box and the true box;

[0034] The calculation formulas for α and v are as follows:

[0035]

[0036] Among them, w gt 、h gt Represents the width and height of the real box, w and h represent the width and height of the predicted box.

[0037] In the above step S12, the virtual data set is created using Unreal 5 software, the virtual data set and the real data set are labeled, and the labeled labels are saved in YOLO format to obtain the power transmission line bird nest data set. In this step, the data is labeled using labelme software.

[0038] In the above step S13, the YOLOv7 network is first trained using a virtual data set. After the training is completed, the real data set is further trained using transfer learning to complete the training of the final model.

[0039] The transfer learning in this step is mainly based on the migration of shared parameters, which is to use the pre-trained model as the initial model and then fine-tune it on the new task. The pre-trained model is usually trained on a large-scale data set, such as a virtual data set used in the present invention. Through the feature representation learned on large-scale data, the model can better capture common features and can adapt to new data more quickly when fine-tuning on a new task, such as the real data set used in the present invention.

[0040] Through the above step S1, the improvement and training of the YOLOv7 network can be completed to achieve the function of detecting whether there is a bird nest in the captured image.

[0041] In the above step S2, the trained model is mainly deployed to the recognition device, and the installation of the entire drone, recognition device, battery, and camera is completed.

[0042] The above step S2 mainly includes the following sub-steps:

[0043] S21 deploys the trained network on the recognition device;

[0044] S22 first completes the assembly of the drone, and then installs the identification equipment, camera, and battery on the drone.

[0045] In the above step S21, the improved YOLOv7 network built by the pytorch deep learning network framework is first converted into the ONNX intermediate conversion format, and then the model is deployed to the hardware recognition device using the deployment framework.

[0046] In the above step S22, the drone is first installed, and then the identification device, camera, battery, etc. are fixed to the drone. The drone, identification device, and camera are conventional devices disclosed in the art. The battery is electrically connected to the identification device and the camera to power the identification device and the camera.

[0047] Through the above S2 step, the YOLOv7 network can be deployed on the recognition device and the installation of the entire drone automatic detection module can be completed.

[0048] In the above step S3, the image of the transmission line tower is captured by controlling the drone and uploaded to the recognition device for bird nest detection. The image captured in this step can be either for local detection or global detection.

[0049] The above step S3 includes the following sub-steps:

[0050] S31 controls the drone equipped with identification equipment and camera to run near the transmission line tower;

[0051] S32 controls the drone to fly around the transmission line tower and take images of the tower from various angles;

[0052] S33 uploads the captured image to the recognition device and completes the detection of the bird's nest.

[0053] In the above step S31, a drone equipped with identification equipment and a camera is controlled to fly near a transmission line tower.

[0054] In the above step S32, the drone is controlled to fly to each designated point, and the camera is used to capture images of the tower at various angles, which may include partial images of the tower or images of the entire tower.

[0055] In the above step S33, the image captured by the camera is uploaded to the recognition device, and the YOLOv7 network is run to detect whether there is a bird nest.

[0056] Through the above step S3, the tower image can be collected and the bird's nest can be detected.

[0057] In the above step S4, after the recognition device completes the bird nest detection, it uploads the processed results to the cloud, and the terminal device displays the results after receiving them. In the specific implementation method, the processed detection results and the original image are transmitted to the local computer. And the results are displayed on the designed software interface.

[0058] Compared with the prior art, the power transmission line foreign body intrusion defect detection method based on improved YOLOv7 provided by the present invention has the following beneficial effects:

[0059] 1. Aiming at the problem that foreign objects invading bird nests on power transmission lines account for a small proportion, are severely blocked, have low recognition accuracy, and have many missed detections, the present invention proposes an improved YOLOv7 network to increase the network's ability to extract key features, and uses kmeans clustering to adjust the proportion of anchor frames to better adapt to the field of foreign object intrusion on power transmission lines;

[0060] 2. The present invention improves the YOLOv7 network by adding a CBAM attention mechanism after the original Conv layer, which can improve the ability of the YOLOv7 network to selectively focus on and process information according to different parts of the input, thereby improving the recognition effect;

[0061] 3. The present invention improves the YOLOv7 network and introduces a clockwise four-scale fusion module. By using convolution kernels of different sizes to extract features, the problem of bird nests being too large or too small due to changes in the shooting position of the drone can be effectively addressed;

[0062] 4. Aiming at the problem that the data set of the power transmission line tower is small, the present invention uses Unreal 5 software to create a virtual data set, and performs data enhancement on the actually collected images, thereby expanding the number of data sets;

[0063] 5. By using drones equipped with image acquisition and bird nest detection modules, intelligent identification of bird nests on transmission line towers is achieved, effectively solving the problems of low efficiency and difficulty in manual detection of bird nests. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1A schematic flow chart of a method for detecting foreign body intrusion defects in power transmission lines based on an improved YOLOv7 network provided in Example 1 of the present invention;

[0065] Figure 2 is a schematic flow chart of step S1 in Example 1;

[0066] Figure 3 This is a schematic diagram of the improved YOLOv7 network structure in Example 1;

[0067] Figure 4 Virtual data samples made with Unreal 5;

[0068] Figure 5 Schematic diagram of the process of step S2 in Example 1;

[0069] Figure 6 This is a physical picture of the drone collection equipment in Example 1;

[0070] Figure 7 is a schematic flow chart of step S3 in Example 1;

[0071] Figure 8 Collecting images of the iron tower in Embodiment 1 of the present invention;

[0072] Fig. 9 For the present invention Figure 8 Bird nest detection result diagram in ;

[0073] Fig.10 Schematic diagram of the improved YOLOv7 network structure in Example 2 of the present invention;

[0074] Fig.11 This is a schematic diagram of the structure of a clockwise four-scale fusion module in Example 2 of the present invention;

[0075] Fig.12 Schematic diagram of dilated convolution in Example 2 of the present invention;

[0076] Fig.13 Collecting images of the iron tower in Embodiment 2 of the present invention;

[0077] Fig.14 For the present invention Fig.13 Bird nest detection result diagram in ;

[0078] Fig.15 This is a diagram of the bird nest detection interface in Example 2 of the present invention. DETAILED DESCRIPTION

[0079] The following examples of implementation of the present invention are given in conjunction with the accompanying drawings, and the technical solution of the present invention is further elaborated and described in detail through the examples. The following examples are only some of the examples of implementation of the present invention. Based on the content of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.

[0080] Example 1

[0081] The method for detecting foreign body intrusion defects in power transmission lines based on the improved YOLOv7 network provided in this embodiment is as follows: Figure 1 As shown, it includes the following steps:

[0082] S1 uses the power transmission line bird nest dataset for transfer learning to train the improved YOLOv7 network;

[0083] S2 completes the deployment of the trained YOLOv7 network and the installation of drones, recognition equipment, cameras, and batteries;

[0084] S3 completes the shooting of on-site tower images to detect bird nests;

[0085] S4 uploads the processed results to the cloud and displays the results.

[0086] The above steps S1-S4 are explained in detail below.

[0087] In the above step S1, in order to solve the problems of small proportion of bird nests and severe occlusion, an improved YOLOv7 network is constructed to improve the accuracy of the model in identifying bird nests. First, the improved YOLOv7 network is constructed, and then the constructed improved YOLOv7 network is trained using the virtual data set and the real data set respectively.

[0088] Specifically, Figure 2 As shown, step S1 includes the following sub-steps:

[0089] S11 builds an improved YOLOv7 network.

[0090] The improved YOLOv7 network provided in this embodiment is as follows: Figure 3 As shown, the YOLOv7 network includes a backbone network for feature extraction and a head network for target recognition.

[0091] The backbone network has an inverted pyramid structure, including an input layer, a first feature extraction unit, a second feature extraction unit, a third feature extraction unit, and a fourth feature extraction unit arranged in sequence. The first feature extraction unit, the second feature extraction unit, and the third feature extraction unit are output to the head network; the features extracted by the fourth feature extraction unit are processed by the pooling unit and then input to the head network.

[0092] The above-mentioned input layer includes three convolution modules arranged in sequence, namely a first convolution module, a second convolution module and a third convolution module.

[0093] The above-mentioned first feature extraction unit, second feature extraction unit, third feature extraction unit and fourth feature extraction unit have the same structure. The first feature extraction unit includes a fourth convolution module and a first multi-branch stacking module arranged in sequence. The second feature extraction unit includes a fifth convolution module and a second multi-branch stacking module arranged in sequence. The third feature extraction unit includes a sixth convolution module and a third multi-branch stacking module arranged in sequence. The fourth feature extraction unit includes a seventh convolution module and a fourth multi-branch stacking module arranged in sequence.

[0094] The above pooling unit uses the SPPCSPC module. The SPPCSPC module includes three eighth convolution modules, three ninth convolution modules, three two-dimensional convolution layers of different scales (5×5, 9×9, 13×13), two tenth convolution modules and an eleventh convolution module. The features input to the SPPCSPC module are respectively input to three eighth convolution modules and three ninth convolution modules. The output features of the last eighth convolution module are concatenated with the features after convolution processing of three two-dimensional convolution layers of different scales and then input to two tenth convolution modules. The outputs of the last ninth convolution module and the last tenth convolution module are concatenated and input to the eleventh convolution module to obtain the output features of the SPPCSPC module.

[0095] The head network processes the input features through path aggregation to obtain several layers of output. In the head network, a CBAM attention mechanism is added after the original Conv layer. The CBAM attention mechanism is a hybrid attention mechanism that combines spatial attention and channel attention mechanisms, which can improve the ability of the YOLOv7 network to selectively focus on and process information according to different parts of the input, thereby improving the recognition effect. In this embodiment, the head network includes a first fusion layer, a second fusion layer, a third fusion layer, a first output layer, a second output layer, and a third output layer.

[0096] In the first fusion layer, the output features of the pooling unit are processed through the first convolutional layer (Conv layer), the first CBAM module and the first upsampling module to obtain the first feature. The output features of the third feature extraction unit are processed through the second convolutional layer (Conv layer) and the second CBAM module to obtain the second feature. The first feature and the second feature are concatenated and then processed through the fifth multi-branch stacking module to obtain the first fusion feature.

[0097] In the second fusion layer, the first fused feature is processed through the first two-dimensional convolution layer and the second upsampling module to obtain the third feature, the output feature of the second feature extraction unit is processed through the third convolution layer (Conv layer) and the third CBAM module to obtain the fourth feature, and the third feature and the fourth feature are concatenated and then processed through the sixth multi-branch stacking module to obtain the second fused feature.

[0098] In the third fusion layer, the second fused feature is processed through the second two-dimensional convolution layer and the third upsampling module to obtain the fifth feature, the output feature of the first feature extraction unit is processed through the fourth convolution layer (Conv layer) and the fourth CBAM module to obtain the sixth feature, and the fifth feature and the sixth feature are concatenated and then processed through the seventh multi-branch stacking module to obtain the third fused feature.

[0099] In the first output layer, the third fusion feature is spliced ​​with the output feature of the first transition module and the second fusion feature, and then processed by the eighth multi-branch stacking module and the first parameterized convolution (Repconv) layer to obtain the first output feature.

[0100] In the second output layer, the first output feature is spliced ​​with the output feature of the second transition module and the first fusion feature, and then processed by the ninth multi-branch stacking module and the second parameterized convolution (Repconv) layer to obtain the second output feature.

[0101] In the third output layer, the second output feature is concatenated with the output feature of the third transition module and the output feature of the pooling unit, and then processed by the tenth multi-branch stacking module and the third parameterized convolution (Repconv) layer to obtain the third output feature.

[0102] The first output feature, the second output feature and the third output feature with better recognition effect are taken as the final recognition result. In this embodiment, the final recognition result is selected by conventional non-maximum suppression (NMS).

[0103] The first convolution module, the second convolution module, the third convolution module, the fourth convolution module, the fifth convolution module, the sixth convolution module, the seventh convolution module, the eighth convolution module, the ninth convolution module, the tenth convolution module and the eleventh convolution module have the same structure, and all include a two-dimensional convolution layer (Conv2D), a BN layer (Batch Normalization layer) and a SiLU activation function.

[0104] The above-mentioned first multi-branch stacking module, second multi-branch stacking module, third multi-branch stacking module, fourth multi-branch stacking module, fifth multi-branch stacking module, sixth multi-branch stacking module, seventh multi-branch stacking module, eighth multi-branch stacking module, ninth multi-branch stacking module and tenth multi-branch stacking module have the same structure, and all adopt the multi-branch stacking module (Multi_Concat_Block) structure in the traditional YOLOv7 network. For the specific structure, see YOLOv7: TrainableBag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors.

[0105] The first CBAM module, the second CBAM module, the third CBAM module and the fourth CBAM module have the same structure. For the specific structure, see CBAM: Convolutional Block Attention Module.

[0106] The first transition module, the second transition module, and the third transition module have the same structure and all adopt the transition module (Transition_Block) structure in the traditional YOLOv7 network. For the specific structure, see YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors.

[0107] The first parameterized convolution (Repconv) layer, the second parameterized convolution (Repconv) layer, and the third parameterized convolution (Repconv) layer have the same structure, and all use the re-parameterized convolution (Repconv) layer structure in the traditional YOLOv7 network. For the specific structure, see RepVGG: Making VGG-style ConvNets Great Again.

[0108] Moreover, in this embodiment, the kmeans clustering algorithm is used to determine the most suitable anchor frame ratio (i.e., aspect ratio) in the YOLOv7 network based on the width and height distribution of the bird nest target frame (i.e., the bounding box of the bird nest in the image) in the collected real data set containing the bird nest, that is, a cluster center is determined by the kmeans clustering algorithm, and the aspect ratio of the anchor frame output by the YOLOv7 network is adjusted according to the anchor frame ratio of the cluster center, so that the aspect ratio of the output anchor frame deviates from the aspect ratio of the anchor frame of the cluster center within 10% (the anchor frame outside this range is directly deleted), thereby optimizing the detection ability of the YOLOv7 network for the bird nest target.

[0109] The loss function used in the above improved YOLOv7 network is CIOU.

[0110] The calculation formula of CIOU is as follows:

[0111]

[0112] Among them, IOU is the intersection area of ​​the predicted box and the real box divided by the area of ​​the two together; ρ(b,b gt ) represents the Euclidean distance between the center points of the predicted box and the true box, and c represents the diagonal distance of the minimum closure area that can contain both the predicted box and the true box;

[0113] The calculation formulas for α and v are as follows:

[0114]

[0115] Among them, w gt 、h gt Represents the width and height of the real box, w and h represent the width and height of the predicted box.

[0116] S12 uses Unreal 5 software to create a virtual data set, annotates the virtual data set and the real data set, and saves the annotated labels in YOLO format.

[0117] In this embodiment, virtual data samples are produced using Unreal 5 software based on the collected real transmission line routes containing bird nests. Based on various conditions such as different weather (such as sunny days, rainy days, foggy days), lighting (changes between day and night), angles (different shooting angles), etc., Unreal 5 software is used to simulate these conditions to generate corresponding virtual data samples, such as Figure 4 All the created virtual data samples constitute a virtual data set.

[0118] The virtual data set and the real data set are annotated with bird nests, and the annotated labels are saved in YOLO format to obtain the power transmission line bird nest data set. In this step, labelme software is used to annotate the data.

[0119] S13 first uses a virtual data set to train the YOLOv7 network. After the training is completed, transfer learning is used to continue training the real data set to complete the training of the final model.

[0120] First, use the virtual data set to train the YOLOv7 network. After the training is completed, use transfer learning to continue training the real data set to complete the training of the final model.

[0121] The transfer learning in this step is mainly based on the migration of shared parameters. It uses the pre-trained model as the initial model and then fine-tunes it on the new task. The pre-trained model is usually trained on a virtual dataset. The training process is based on the loss function given above and the traditional YOLOv7 network training method. Then it is fine-tuned on the real dataset to adapt to new data more quickly.

[0122] The above step S2 mainly deploys the YOLOv7 network trained in step S1 to the recognition device and completes the installation of the entire drone, recognition device, battery, and camera.

[0123] Specifically, Figure 5 As shown, step S2 includes the following sub-steps:

[0124] S21 deploys the trained network on the recognition device.

[0125] In this step, the YOLOv7 network built by the pytorch deep learning network framework is first converted into the ONNX intermediate conversion format, and then the model is deployed to the hardware recognition device using the deployment framework. In this embodiment, the recognition device is NVIDIA Jetson Orin Nano.

[0126] S22 first completes the assembly of the drone, and then installs the identification equipment, camera, and battery on the drone.

[0127] In this step, the drone is first installed, and then the identification device, camera, battery, etc. are fixed to the drone to obtain the drone data collection device. The drone, identification device and camera use conventional devices that have been disclosed in the field. The battery is electrically connected to the identification device and the camera to power the identification device and the camera. The actual installation picture of the drone data collection device finally built in this step is as follows: Figure 6 shown.

[0128] In the above step S3, the camera of the drone is controlled to capture images of the transmission line tower and the images are uploaded to the recognition device for bird nest detection.

[0129] Specifically, Figure 7 As shown, step S3 includes the following sub-steps:

[0130] S31 controls the drone equipped with identification equipment and camera to run near the transmission line tower.

[0131] In this step, the drone equipped with the identification device and the camera is controlled to fly near the transmission line tower, that is, a location capable of collecting images of part or all of the transmission line tower.

[0132] S32 controls the drone to fly around the transmission line tower and take images of the tower from various angles.

[0133] In this step, the drone is controlled to fly around each layer of the transmission line tower and take images of the tower from various angles. The images of the transmission line tower taken are as follows: Figure 8 shown.

[0134] S33 uploads the captured image to the recognition device and completes the detection of the bird's nest.

[0135] In this step, the collected image of the transmission line tower is transmitted back to the recognition device, and the improved YOLOv7 network is run to detect whether there is a bird nest in the image and select it. Figure 8 The test results are as follows Fig. 9 shown.

[0136] In the above step S4, the detection result of step S3 is uploaded to the terminal device.

[0137] In this embodiment, after the recognition device completes the bird nest detection, it uploads the processed results to the cloud, and then transmits them from the cloud to the terminal device (here, the local computer), and the terminal device displays the results after receiving them.

[0138] Example 2

[0139] The method for detecting foreign body intrusion defects in a power transmission line based on an improved YOLOv7 network provided in this embodiment includes the following steps:

[0140] S1 uses the power transmission line bird nest dataset for transfer learning to train the improved YOLOv7 network;

[0141] S2 completes the deployment of the trained YOLOv7 network and the installation of drones, recognition equipment, cameras, and batteries;

[0142] S3 completes the capture of on-site tower images for bird nest detection;

[0143] S4 uploads the processed results to the cloud and displays the results.

[0144] The explanation of the above steps S1 to S4 is basically the same as that in Example 1, except that the improved YOLOv7 network used is different.

[0145] like Fig.10 As shown, the improved YOLOv7 network provided in this embodiment includes, in addition to the backbone network and the head network, a clockwise quad-scale fusion module (CQSFM) set before the input layer. Fig.11As shown in the figure, the clockwise four-scale fusion module includes four different scales of two-dimensional convolution layers, a maximum pooling layer and a splicing unit; first, the input image is convolved with four different scales of two-dimensional convolution layers (in this embodiment, the convolution kernel sizes are 1×1, 3×3, 5×5, and 7×7, respectively), and then the size is reduced to half of the original size through a 2×2 maximum pooling layer; then, the four convolution outputs are spliced ​​in a clockwise direction to form a new image as the input of the input layer. By using convolution kernels of different sizes to extract features, the problem of bird nests being too large or too small due to changes in the shooting position of the drone can be effectively addressed.

[0146] Moreover, in this embodiment, all two-dimensional convolutional layers in the improved YOLOv7 network are replaced with hole convolutional layers. Fig.12 As shown, the expansion coefficient of the dilated convolution layer in this embodiment is 2. The dilated convolution introduces intervals between the convolution kernels, thereby expanding the receptive field and capturing a wider range of contextual information without increasing the amount of computation.

[0147] The constructed improved YOLOv7 network is trained through step S1; then the trained improved YOLOv7 network is deployed on the recognition device according to step S2, and the drone, recognition device, camera and battery are installed; then the image of the transmission line tower to be recognized is collected according to step S3, such as Fig.13 As shown; the collected image of the transmission line tower is transmitted back to the recognition device, and the improved YOLOv7 network is run to obtain the detection result, as shown Fig.14 As shown; finally, through step S4, after the identification device completes the bird nest detection, the processed results are uploaded to the cloud, and then transmitted from the cloud to the terminal device (here is the local computer), and the terminal device displays the results after receiving them.

[0148] In this embodiment, the terminal device also displays the received detection results on the designed software interface (here, the bird nest detection interface), such as Fig.15 As shown in the figure, it can be seen that from the bird nest detection interface, information such as the image containing the bird nest, the bird nest anchor frame, and the location of the bird nest on the transmission line tower can be given.

[0149] Those skilled in the art will appreciate that the embodiments herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A method for detecting foreign body intrusion defects in power transmission lines based on an improved YOLOv7 network, characterized in that: The main steps include: S1 uses the power transmission line bird nest dataset for transfer learning to train the improved YOLOv7 network; step S1 includes the following sub-steps: S11 builds an improved YOLOv7 network; uses the kmeans clustering algorithm to adjust the proportion of anchor boxes in the YOLOv7 network; S12 used Unreal 5 software to create a virtual data set, labeled the virtual data set and the real data set, and saved the labeled labels in yolo format; S13 first uses a virtual data set to train the YOLOv7 network. After the training is completed, transfer learning is used to continue training the real data set to complete the training of the final model. S2 completes the deployment of the trained YOLOv7 network and the installation of drones, recognition equipment, cameras, and batteries; S3 completes the shooting of on-site tower images to detect bird nests; S4 uploads the processed results to the cloud and displays the results.

2. According to the method for detecting foreign body intrusion defects in power transmission lines based on the improved YOLOv7 network in claim 1, it is characterized in that: In step S11, according to the width and height distribution of the bird nest target box in the collected real data set containing the bird nest, the kmeans clustering algorithm is used to determine the anchor box ratio in the YOLOv7 network.

3. The method for detecting foreign body intrusion defects in power transmission lines based on the improved YOLOv7 network according to claim 1 is characterized in that: The feature map output by the first feature extraction unit in the backbone network of the YOLOv7 network is used in subsequent operations.

4. The method for detecting foreign body intrusion defects in power transmission lines based on the improved YOLOv7 network according to claim 3 is characterized in that: The YOLOv7 network includes a backbone network for feature extraction and a head network for target recognition; the backbone network has an inverted pyramid structure, including an input layer, a first feature extraction unit, a second feature extraction unit, a third feature extraction unit and a fourth feature extraction unit arranged in sequence; The first feature extraction unit, the second feature extraction unit and the third feature extraction unit are output to the head network; the features extracted by the fourth feature extraction unit are processed by the pooling unit and then input to the head network; the head network processes the input features by path aggregation to obtain several layers of output; the input layer includes three convolution modules arranged in sequence; the first feature extraction unit, the second feature extraction unit, the third feature extraction unit and the fourth feature extraction unit have the same structure, including convolution modules and multi-branch stacking modules arranged in sequence; the pooling unit uses the SPPCSPC module; the convolution module includes a two-dimensional convolution layer, a BN layer and a SiLU activation function.

5. The method for detecting foreign body intrusion defects in power transmission lines based on the improved YOLOv7 network according to claim 4 is characterized in that: A clockwise four-scale fusion module is also provided before the input layer, and the clockwise four-scale fusion module includes four two-dimensional convolutional layers of different scales, a maximum pooling layer and a splicing unit; first, the input image is convolved with four two-dimensional convolutional layers of different scales, and then passed through the maximum pooling layer to reduce the size to half of the original size; then, the four convolution outputs are spliced ​​in a clockwise direction to form a new image as the input of the input layer.

6. The method for detecting foreign body intrusion defects in power transmission lines based on an improved YOLOv7 network according to claim 4 or 5, characterized in that: In the head network, the CBAM attention mechanism is added after the original Conv layer.

7. The method for detecting foreign body intrusion defects in power transmission lines based on the improved YOLOv7 network according to claim 6 is characterized in that: The head network includes a first fusion layer, a second fusion layer, a third fusion layer, a first output layer, a second output layer and a third output layer; In the first fusion layer, the output features of the pooling unit are processed by the convolution layer, the CBAM module and the upsampling module to obtain the first feature, the output features of the third feature extraction unit are processed by the convolution layer and the CBAM module to obtain the second feature, and the first feature and the second feature are concatenated and then processed by the multi-branch stacking module to obtain the first fusion feature; In the second fusion layer, the first fusion feature is processed through a two-dimensional convolution layer and an upsampling module to obtain a third feature, the output feature of the second feature extraction unit is processed through a convolution layer and a CBAM module to obtain a fourth feature, and the third feature and the fourth feature are concatenated and then processed through a multi-branch stacking module to obtain a second fusion feature; In the third fusion layer, the second fusion feature is processed by a two-dimensional convolution layer and an upsampling module to obtain a fifth feature, the output feature of the first feature extraction unit is processed by a convolution layer and a CBAM module to obtain a sixth feature, and the fifth feature and the sixth feature are concatenated and then processed by a multi-branch stacking module to obtain a third fusion feature; In the first output layer, the output feature of the third fusion feature after the transition module is concatenated with the second fusion feature and then processed by the multi-branch stacking module and the re-parameterized convolution layer to obtain the first output feature; In the second output layer, the first output feature is spliced ​​with the output feature of the transition module and the first fusion feature, and then processed by the multi-branch stacking module and the re-parameterized convolution layer to obtain the second output feature; In the third output layer, the second output feature is concatenated with the output feature of the transition module and the output feature of the pooling unit, and then processed by a multi-branch stacking module and a re-parameterized convolutional layer to obtain the third output feature.

8. The method for detecting foreign body intrusion defects in power transmission lines based on the improved YOLOv7 network according to claim 6 is characterized in that: Improve the two-dimensional convolutional layer in the YOLOv7 network and replace it with a hole convolutional layer.

9. The method for detecting foreign body intrusion defects in power transmission lines based on the improved YOLOv7 network according to claim 6 is characterized in that: Improve the YOLOv7 network and change the loss function to CIOU; the calculation formula of CIOU is as follows: Among them, IOU is the intersection area of ​​the predicted box and the real box divided by the area of ​​the two together; ρ(b,b gt ) represents the Euclidean distance between the center points of the predicted box and the true box, and c represents the diagonal distance of the minimum closure area that can contain both the predicted box and the true box; The calculation formulas for α and v are as follows: where w gt 、h gt Represents the width and height of the real box, w and h represent the width and height of the predicted box.

10. The method for detecting foreign body intrusion defects in power transmission lines based on the improved YOLOv7 network according to claim 1, characterized in that: The step S2 mainly includes the following sub-steps: S21 deploys the trained network on the recognition device; S22 first completes the assembly of the drone, and then installs the identification device, camera, and battery on the drone; The step S3 comprises the following sub-steps: S31 controls the drone equipped with identification equipment and camera to run near the transmission line tower; S32 controls the drone to fly around the transmission line tower and take images of the tower from various angles; S33 uploads the captured image to the recognition device and completes the detection of the bird's nest.