Grid line-crossing bird recognition method and system based on multi-scale feature enhancement
The RBOD model, leveraging multi-scale features and a lightweight super-resolution model, addresses the inefficiencies of traditional bird collision prevention methods by accurately identifying rare birds near power lines, enhancing detection capabilities and supporting power system safety and conservation.
Patent Information
- Application Number
- CN202510618475.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-14
AI Technical Summary
Traditional bird-proof measures such as bird stabbing and bird repellers have passive protection and ecological interference, manual inspection efficiency is low, and it is difficult to fully cover the risk of bird impacting transmission lines. It is difficult to achieve efficient and accurate monitoring and identification of bird activities in the existing technology.
Based on the YOLOv10 model, the RBOD rare bird object detection model is improved, and through multi-scale feature enhancement feature extraction network, fusion network and detection network, combined with lightweight image super-resolution reconstruction model, a rare bird image database is constructed to achieve efficient identification of rare birds.
It improves the accuracy and speed of rare bird identification, provides technical support for the safe operation of the power system and the protection of rare birds, reduces bird impact events, and balances ecological protection and industrial development.
Smart Images

Figure CN120125920B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of transmission line inspection, and particularly relates to a power grid wire-striking bird recognition method and system based on multi-scale feature enhancement. Background Art
[0002] Birds hitting transmission lines can cause accidents such as short circuits, tripping, and even fires of transmission lines, resulting in huge economic losses. Traditional solutions such as physical devices like anti-bird spines and bird repellers can partially alleviate the problem, but they have limitations such as passive protection and interference with the ecological balance. Manual inspection is limited by efficiency and safety and is difficult to achieve comprehensive coverage. In this context, the rise of artificial intelligence technology provides a new path for the collaborative optimization of the power system and bird protection.
[0003] In recent years, the application of artificial intelligence in bird protection in the power system has become increasingly widespread. Based on advanced artificial intelligence algorithms, operation and maintenance personnel can achieve accurate monitoring of bird activities, and through high-precision and high-efficiency technical means, they can identify and locate bird activities near transmission lines in real time, thereby providing data support for subsequent protection measures. The intelligent system can achieve efficient recognition and location of bird activities near transmission lines, providing strong technical support for the safe operation of the power system and bird protection. The wide application of this technology can not only reduce bird strike incidents but also provide a scientific basis for the balance between ecological protection and industrial development. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a power grid wire-striking bird recognition method and system based on multi-scale feature enhancement. The present invention improves on the basis of the YOLOv10 model to construct an RBOD rare bird target detection model, which can accurately identify power grid wire-striking birds and provide strong technical support for the safe operation of the power system and the protection of rare birds.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A power grid wire-striking bird recognition method based on multi-scale feature enhancement, which constructs and trains an RBOD rare bird target detection model; uses the trained RBOD rare bird target detection model to recognize power grid wire-striking birds;
[0006] The RBOD rare bird target detection model includes three parts: a feature extraction network, a feature fusion network, and a detection network. The feature extraction network consists of an SRFD module, a first CWDB module, a first DRFD module, a second CWDB module, a second DRFD module, a third CWDB module, a third DRFD module, a fourth CWDB module, an SPPF module, and a PSA module in sequence; the feature fusion network is used to effectively integrate features of different scales; the DRFD module includes an upper branch and a lower branch: the lower branch of the DRFD module is pixel cutting downsampling. In the upper branch of the DRFD module, first, grouped convolution is performed, and then through two branches. The first branch completes depthwise separable convolution and then uses the GELU activation function for feature enhancement. The second branch performs max pooling downsampling; the output features of the two branches are concatenated and convolved with the output of the lower branch, and finally, the output of the DRFD module is obtained;
[0007] The CWDB module consists of a CBS module, a separation module, several WDBB modules, a fusion module, and another CBS module in sequence.
[0008] Further preferably, the processing process of the feature fusion network is as follows: the output feature F3 of the PSA module is upsampled by the first upsampling module and then fused with the output feature F2 of the third CWDB module through the first fusion module to obtain the fused feature C1, which is input to the fifth CWDB module for processing; the output feature C2 of the fifth CWDB module is processed by the upsampling module and then fused with the output feature F1 of the second CWDB module through the second fusion module to obtain the fused feature C3; the fused feature C3 is processed by the sixth CWDB module, and the output feature C4 of the sixth CWDB module is processed by the fourth DRFD module and then fused with the output feature C2 of the fifth CWDB module through the third fusion module to obtain the fused feature C5. The fused feature C5 is input to the seventh CWDB module for processing. The output feature C6 of the seventh CWDB module is processed by the fifth DRFD module and then fused with F3 through the fourth fusion module to obtain the fused feature C7. The fused feature C7 is input to the C2FCIB module for processing. The output feature C4 of the sixth CWDB module, the output feature C6 of the seventh CWDB module, and the output feature C8 of the C2FCIB module are used as the inputs of the three detection modules of the detection network.
[0009] Further preferably, the WDBB module divides the input features into six branches for processing. The first branch is sequentially processed by a 1×1 ordinary convolution and a batch normalization layer; the second branch is sequentially processed by a 1×1 ordinary convolution, the first batch normalization layer, a k×k ordinary convolution, and the second batch normalization layer; the third branch is sequentially processed by a 1×1 ordinary convolution, the first batch normalization layer, an average pooling layer, and the second batch normalization layer; the fourth branch is sequentially processed by a k×k ordinary convolution and a batch normalization layer; the fifth branch is sequentially processed by a 1×k ordinary convolution and a batch normalization layer; the sixth branch is sequentially processed by a k×1 ordinary convolution and a batch normalization layer; the features processed by the six branches are subjected to convolutional summation and output features after passing through the activation function SiLU.
[0010] Further preferably, the operation mode of the PSA module is as follows: the module is divided into seven parts, namely the first ordinary convolution, the multi-head self-attention module, the first feature splicing module, the feed-forward network, the second feature splicing module, the fusion module, and the second ordinary convolution. The features output by the first ordinary convolution are respectively input to the multi-head self-attention module, the first feature splicing module, and the fusion module. The features after splicing by the first feature splicing module are respectively input to the feed-forward network and the second feature splicing module. The features are fused by the fusion module and then input to the second ordinary convolution for convolutional processing to obtain the final output of the PSA module.
[0011] Further preferably, the processing process of the C2FCIB module is as follows: the image features are input to a CBS module for convolutional processing, then processed by a separation module, and then sequentially processed by several CIB modules. Finally, the features of the separation module and several CIB modules are subjected to feature fusion through a fusion module, and the fused features are input to another CBS module for processing to obtain the output of the C2FCIB module.
[0012] Further preferably, the shallow feature extraction module is the SRFD module. The SRFD module includes three stages. The first stage uses ordinary convolution. The features output in the first stage enter the second stage. The upper branch of the second stage consists of grouped convolution and depthwise separable convolution, and the lower branch of the second stage is pixel cutting downsampling. The features of the upper branch and the lower branch are subjected to splicing and convolutional operations and then enter the third stage. In the upper branch of the third stage, grouped convolution is first performed, and then depthwise separable convolution and max-pooling downsampling are respectively performed through two branches. The features output by the depthwise separable convolution and max-pooling downsampling are spliced and convolved with the output of the lower branch of the third stage, and finally the output features of the SRFD module are obtained.
[0013] Further preferably, the detection module consists of two branches. The operation mode of the first branch is as follows: after the input features are convolved by the CBS module, they pass through the MultiSEAM module and ordinary convolution in sequence to obtain the prediction loss. The operation mode of the second branch is as follows: after the input features are processed by two consecutive CBS modules, they pass through the MultiSEAM module and ordinary convolution in sequence to obtain the classification loss.
[0014] Further preferably, according to the historical data of rare bird wire-striking accidents in the power grid, rare bird images are collected through transmission line monitoring cameras and network resources. The lightweight image super-resolution reconstruction model is used to super-resolve the rare bird images to improve the image resolution. After the image reconstruction is completed, a dataset of rare birds that accidentally strike the transmission line is constructed. The dataset of rare birds that accidentally strike the transmission line is used to train the RBOD rare bird target detection model.
[0015] The present invention also provides a power grid wire-striking bird recognition system based on multi-scale feature enhancement, including an image acquisition device and a power grid wire-striking bird recognition device. The image acquisition device is used to collect bird images within a certain distance range of the transmission line. The power grid wire-striking bird recognition device is built-in with the RBOD rare bird target detection model for bird recognition.
[0016] The present invention also provides an electronic device, including a memory and a processor. The memory stores computer-readable instructions. When the instructions are executed by the processor, the processor is caused to implement the above-mentioned power grid wire-striking bird recognition method based on multi-scale feature enhancement.
[0017] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned power grid wire-striking bird recognition method based on multi-scale feature enhancement is implemented.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: First, rare bird images are collected through power transmission line monitoring cameras and network resources, and a lightweight image super-resolution reconstruction model is used to perform super-resolution reconstruction on the rare bird images, aiming to improve the image resolution, which can improve the effectiveness of the images for the model. After image enhancement, a database of rare birds hitting the power grid line images is constructed. Second, an RBOD rare bird target detection model is constructed, and the specific task of rare bird recognition is improved through three modules: First, for the feature extraction network, the SRFD module is first used for shallow feature extraction, which retains the original feature information and fuses local information, effectively reducing the loss of feature information in the initial stage. Subsequently, the DRFD module is used as the deep feature extraction, and the model captures feature information of different scales through multi-scale fusion and reduces the model calculation amount. The use of the SRFD module and the DRFD module in the model can effectively improve the feature extraction ability of the model for bird images. Second, the WDBB module is used to improve the feature extraction network and the feature fusion network. With its multi-branch structure, the WDBB module enables the model to learn diverse feature representations and perform feature enhancement on them, effectively enhancing the detection ability of the model for rare birds. Third, the MultiSEAM module is used to improve the detection module. The MultiSEAM module can accurately extract bird features in complex scenarios and enhance the feature representation ability of the detection module. After the model construction, the training set and validation set images are used to train the model, and finally the trained model is used to detect the test set images. The results show that the RBOD rare bird target detection model proposed by the present invention can accurately and quickly identify rare birds appearing around the line, and can provide technical support for power grid operation and maintenance personnel in bird recognition and protection. Description of the Drawings
[0019] Figure 1 It is a flowchart of the method of the present invention.
[0020] Figure 2 It is a flowchart of the lightweight image super-resolution reconstruction model.
[0021] Figure 3 It is a schematic diagram of the RBOD rare bird target detection model.
[0022] Figure 4 It is a schematic diagram of the SRFD module.
[0023] Figure 5 It is a schematic diagram of the DRFD module.
[0024] Figure 6 It is a schematic diagram of the CWDB module.
[0025] Figure 7 It is a schematic diagram of the CBS module.
[0026] Figure 8 It is a schematic diagram of the DWBB module.
[0027] Figure 9 It is a schematic diagram of the PSA module.
[0028] Figure 10 It is a schematic diagram of the C2FCIB module.
[0029] Figure 11 It is a schematic diagram of the CIB module.
[0030] Figure 12 It is a schematic diagram of the detection module.
[0031] Figure 13 It is a schematic diagram of the MultiSEAM module.
[0032] Figure 14 It is a schematic diagram of the CSM module.
[0033] Figure 15 It is a schematic diagram of the residual branch. Specific implementation manners
[0034] The following further describes the present invention in conjunction with embodiments. It is necessary to point out here that the following embodiments are only used to further illustrate the present invention and should not be construed as limiting the protection scope of the present invention. Some non-essential improvements and adjustments made by those skilled in the art according to the above invention content still fall within the protection scope of the present invention.
[0035] As Figure 1 shown, the power grid wire-striking bird recognition method based on multi-scale feature enhancement includes the following steps:
[0036] S1. Construct a rare bird image dataset for mis-collision with transmission lines: According to the historical data of rare bird wire-striking accidents in the power grid, collect rare bird images through transmission line monitoring cameras and network resources. Since the transmission line monitoring cameras have the phenomenon of low resolution and unclear bird images, a lightweight image super-resolution reconstruction model is used to super-resolve the rare bird images to improve the image resolution. After completing the image reconstruction, construct a rare bird image dataset for mis-collision with transmission lines;
[0037] S2. Improve based on the YOLOv10 model to construct an RBOD rare bird target detection model;
[0038] S3. Train the RBOD rare bird target detection model: First, use the labelimg image annotation tool to annotate the rare bird image dataset of birds hitting transmission lines constructed in step S1, and divide the rare bird image dataset of birds hitting transmission lines into a training set, a validation set, and a test set. Adopt the idea of transfer learning, load the pre-trained weights that have been trained, and use the divided training set and validation set to train the RBOD rare bird target detection model;
[0039] S4. Use the trained RBOD rare bird target detection model to identify birds hitting the power grid.
[0040] In step S1 of this embodiment, eleven species of birds, namely black-necked cranes, saker falcons, grey cranes, oriental white storks, great bustards, white-naped cranes, little egrets, white spoonbills, common kestrels, whooper swans, and northern goshawks, are selected as the identification objects. There are two hundred images of each bird species, a total of two thousand two hundred rare bird species images as the original dataset. Use a lightweight image super-resolution reconstruction model to expand the original dataset. Each bird species image is expanded to three hundred, a total of three thousand three hundred rare bird species image samples, and this is used as the dataset for the subsequent algorithm. The lightweight image super-resolution reconstruction model is as Figure 2 shown. First, input the image, and then perform image cropping, image flipping, and feature transformation. Train the model for the teacher model and the student model, perform knowledge distillation, and use the trained student model for image super-resolution reconstruction.
[0041] In this embodiment, first, annotate the bird images using labelimg according to the scientific names corresponding to the eleven bird species to generate the.txt files required for model training. After the annotation is completed, divide the pictures and labels into a training set, a validation set, and a test set according to the ratio of 8:1:1. Load the pre-trained weights of the model trained based on the COCO open-source large-scale image dataset, and use the training set and the validation set to train the RBOD rare bird target detection model. Set the number of model training rounds to 200, the batch training size to 64, use the SGD optimizer as the optimizer, set the weight decay to 0.0005 to prevent overfitting, set the model confidence to 0.7, and set the initial learning rate to 0.01.
[0042] Specifically, the RBOD rare bird target detection model of this embodiment includes three parts: a feature extraction network, a feature fusion network, and a detection network, as Figure 3As shown, the feature extraction network consists of ten parts in sequence: the first part is the SRFD module, and the second to eighth parts are alternately stacked by four CWDB modules and three DRFD modules (in sequence: the first CWDB module, the first DRFD module, the second CWDB module, the second DRFD module, the third CWDB module, the third DRFD module, and the fourth CWDB module), the ninth part is the SPPF module, and the tenth part is the PSA module; the output feature F1 (large-scale feature) of the second CWDB module, the output feature F2 (medium-scale feature) of the third CWDB module, and the output feature F3 (small-scale feature) of the PSA module are used as the three input features of the feature fusion network.
[0043] The feature fusion network is used to achieve the effective integration of features at different scales. The process is as follows: F3 is upsampled by the first upsampling module and then fused with F2 through the first fusion module to obtain the fused feature C1, which is input to the fifth CWDB module for processing; the output feature C2 of the fifth CWDB module is processed by the upsampling module and then fused with F1 through the second fusion module to obtain the fused feature C3; the fused feature C3 is processed by the sixth CWDB module, and the output feature C4 of the sixth CWDB module is processed by the fourth DRFD module and then fused with the output feature C2 of the fifth CWDB module through the third fusion module to obtain the fused feature C5. The fused feature C5 is input to the seventh CWDB module for processing. The output feature C6 of the seventh CWDB module is processed by the fifth DRFD module and then fused with F3 through the fourth fusion module to obtain the fused feature C7. The fused feature C7 is input to the C2FCIB module for processing. The output feature C4 of the sixth CWDB module, the output feature C6 of the seventh CWDB module, and the output feature C8 of the C2FCIB module are used as the inputs of the three detection modules of the detection network, corresponding to the detections at large, medium, and small scales respectively.
[0044] It should be particularly noted that the numbering of each module in the present invention, such as the first, second, third, etc., is only for the convenience of understanding and explaining the model structure. The structures of the same type of modules with different numbers are the same, but the model parameters may be different, which is easy to understand for those skilled in the art.
[0045] Such as Figure 4As shown in the figure, the SRFD module includes three stages. In the first stage, ordinary convolution is used. The features output from the first stage enter the second stage. The upper branch of the second stage consists of grouped convolution and depthwise separable convolution. The lower branch of the second stage is pixel cutting downsampling. After the features of the upper branch and the lower branch are concatenated and subjected to convolution operations, they enter the third stage. The lower branch of the third stage is pixel cutting downsampling. In the upper branch of the third stage, first grouped convolution is performed, and then depthwise separable convolution and max-pooling downsampling are respectively carried out through two branches. The features output from the depthwise separable convolution and the max-pooling downsampling are concatenated and convolved with the output of the lower branch of the third stage, and finally the output features of the SRFD module are obtained. Through the feature enhancement layer and multi-scale feature fusion, the SRFD module significantly improves the shallow feature extraction ability and can effectively improve the model efficiency for low-resolution and small target scenarios.
[0046] As Figure 5 shown in the figure, the DRFD module includes an upper branch and a lower branch: the lower branch of the DRFD module is pixel cutting downsampling. In the upper branch of the DRFD module, first grouped convolution is performed, and then through two branches, after the first branch completes depthwise separable convolution, the GELU activation function is used for feature enhancement, and the second branch performs max-pooling downsampling to retain key feature information. The output features of the two branches are concatenated and convolved with the output of the lower branch, and finally the output of the DRFD module is obtained. For the fine features after preliminary processing, the DRFD module provides a smoother training gradient through the GELU function, making the model easy to converge and retain more feature information, and improving the model's expression ability for deep features.
[0047] As Figure 6 shown in the figure, the processing process of the CWDB module is as follows: the image features are input into a CBS module for convolution processing, then after being processed by the separation module, they are successively processed by several WDBB modules, and finally the features of the separation module and several WDBB modules are fused through the fusion module. The fused features are input into another CBS module for processing to obtain the output of the CWDB module.
[0048] As Figure 7 shown in the figure, the CBS module consists of ordinary convolution (Conv), batch normalization layer, and activation function in sequence.
[0049] As Figure 8As shown in the figure, the WDBB module divides the input features into six branches for processing. The first branch is processed by a 1×1 ordinary convolution and a batch normalization layer in sequence; the second branch is processed by a 1×1 ordinary convolution, the first batch normalization layer, a k×k ordinary convolution, and the second batch normalization layer in sequence; the third branch is processed by a 1×1 ordinary convolution, the first batch normalization layer, an average pooling layer, and the second batch normalization layer in sequence; the fourth branch is processed by a k×k ordinary convolution and a batch normalization layer in sequence; the fifth branch is processed by a 1×k ordinary convolution and a batch normalization layer in sequence; the sixth branch is processed by a k×1 ordinary convolution and a batch normalization layer in sequence; the features processed by the six branches are convolved and summed, and the output features are obtained after passing through the activation function SiLU. The WDBB module enhances the feature representation, reduces the computational complexity, and improves the robustness and efficiency of the model through the multi-branch structure and the fusion of convolution kernels.
[0050] As Figure 9 shown, the PSA module operates as follows: The module is divided into seven parts, which are the first ordinary convolution, the multi-head self-attention module, the first feature splicing module, the feed-forward network, the second feature splicing module, the fusion module, and the second ordinary convolution in sequence. The output features of the first ordinary convolution are respectively input to the multi-head self-attention module, the first feature splicing module, and the fusion module. The spliced features of the first feature splicing module are respectively input to the feed-forward network and the second feature splicing module. After the feature fusion by the fusion module, it is input to the second ordinary convolution for convolution processing to obtain the final output of the PSA module.
[0051] As Figure 10 shown, the processing process of the C2FCIB module is as follows: The image features are input to a CBS module for convolution processing, then processed by a separation module, and then processed by several CIB modules in sequence. Finally, the features of the separation module and several CIB modules are fused by a fusion module, and the fused features are input to another CBS module for processing to obtain the output of the C2FCIB module.
[0052] As Figure 11 shown, the CIB module is composed of three depthwise separable convolutions and two CBS modules connected in series in sequence. The connection method is the first depthwise separable convolution, the first CBS module, the second depthwise separable convolution, the second CBS module, and the third depthwise separable convolution.
[0053] As Figure 12 shown, the detection module is composed of two branches. The operation mode of the first branch is as follows: After the input features are convolved by the CBS module, they are processed by the MultiSEAM module and an ordinary convolution in sequence to obtain the prediction loss; the operation mode of the second branch is as follows: After the input features are processed by two consecutive CBS modules, they are processed by the MultiSEAM module and an ordinary convolution in sequence to obtain the classification loss.
[0054] As Figure 13 shown, the MultiSEAM module operates as follows: the original input features are respectively input into three CSM modules for processing. The output features obtained by the three CSM modules are then convolutionally summed with the original input features. After summation, the features are respectively subjected to average pooling and fully connected operations. The obtained output is then multiplied by the original input features to obtain the final output features. The MultiSEAM module enables the features to pass through a multi-channel process and then through a fully connected network, fusing the information between channels and enabling the network to completely capture spatial information.
[0055] As Figure 14 shown, the CSM module operates as follows: after the features are input into ordinary convolution, they sequentially pass through the activation function SiLU, the batch normalization layer, and the residual branch processing. The output of the residual branch is added to the features output by the batch normalization layer, and then passes through ordinary convolution, the activation function SiLU, and the batch normalization layer to obtain the output features; as Figure 15 shown, the residual branch is sequentially composed of a depthwise separable convolution, an activation function, and a batch normalization layer. Introducing the residual branch into the network solves the problem of gradient disappearance during training and ensures that shallow features can be directly transmitted to deep layers. The network can focus on learning the residuals of the input features, thereby extracting more discriminative feature information.
[0056] In this embodiment, the RBOD rare bird target detection model that has been trained in step S3 is used to detect the images in the rare bird test set. The deep learning environment used in this embodiment is: a hardware environment with Nvidia GeForce GTX 3060, 12G of video memory, and 16G of memory, and a software environment with Visual Studio Code 2019, CUDA12.1, and CuDNN8.9.6.50. The assembly language used is the python language. According to the confidence level set in step S3, the prediction boxes below the confidence threshold are removed. The model is evaluated using the conventional object detection model evaluation criteria, the mean average precision (mAP) and the detection speed (FPS), and is compared with the YOLOv10 model. The experimental results are as follows.
[0057]
[0058] From the above results, it can be seen that compared with the YOLOv10 model, the model of this patent has obvious improvements in both detection accuracy and detection speed. Moreover, the present invention adopts a lightweight image super-resolution reconstruction model, which can also effectively improve the utilization rate of the photos taken by the transmission line monitoring camera and greatly improve the algorithm recognition efficiency. Deploying the overall model of this patent to the actual scenario can provide technical support for operation and maintenance personnel to identify and protect birds.
[0059] The second embodiment of the present invention provides a power grid wire-striking bird recognition system based on multi-scale feature enhancement, including an image acquisition device and a power grid wire-striking bird recognition device. The image acquisition device is used to collect bird images within a certain distance range of the transmission line. The power grid wire-striking bird recognition device is built-in with an RBOD rare bird target detection model for bird recognition.
[0060] The third embodiment of the present invention provides an electronic device, including a memory and a processor. The memory stores computer-readable instructions. When the instructions are executed by the processor, the processor is enabled to implement the above-mentioned power grid wire-striking bird recognition method based on multi-scale feature enhancement.
[0061] The fourth embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned power grid wire-striking bird recognition method based on multi-scale feature enhancement is implemented.
[0062] The above only expresses the preferred embodiments of the present invention and does not limit the present invention in other forms. Any person skilled in the art may use the disclosed content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for identifying birds hitting power lines based on multi-scale feature enhancement, characterized in that Based on the YOLOv10 model as the basic model, the RBOD rare bird target detection model is constructed and trained; according to the historical data of rare bird line collision accidents on the power grid, rare bird images are collected through transmission line monitoring cameras and network resources, and a lightweight image super-resolution reconstruction model is used to super-resolve the rare bird images to improve the image resolution. After the image reconstruction is completed, a dataset of rare birds that accidentally collide with the transmission line is constructed; the dataset of rare birds that accidentally collide with the transmission line is used to train the RBOD rare bird target detection model; the trained RBOD rare bird target detection model is used to identify the birds that collide with the power grid line; The RBOD rare bird target detection model consists of three parts: a feature extraction network, a feature fusion network, and a detection network. The feature extraction network is sequentially composed of a shallow feature extraction module, a first CWDB module, a first DRFD module, a second CWDB module, a second DRFD module, a third CWDB module, a third DRFD module, a fourth CWDB module, an SPPF module, and a PSA module; the feature fusion network is used to effectively integrate features of different scales; The DRFD module includes an upper branch and a lower branch: the lower branch of the DRFD module is pixel cutting downsampling. In the upper branch of the DRFD module, first, grouped convolution is performed, and then through two branches. The first branch completes depthwise separable convolution and then uses the GELU activation function for feature enhancement. The second branch performs max-pooling downsampling; the output features of the two branches are concatenated and convolved with the output of the lower branch, and finally, the output of the DRFD module is obtained; The CWDB module is sequentially composed of a CBS module, a separation module, several WDBB modules, a fusion module, and another CBS module; the CBS module is sequentially composed of a common convolution, a batch normalization layer, and an activation function; The WDBB module processes the input features in six branches. The first branch is sequentially processed by a 1×1 common convolution and a batch normalization layer; the second branch is sequentially processed by a 1×1 common convolution, a first batch normalization layer, a k×k common convolution, and a second batch normalization layer; the third branch is sequentially processed by a 1×1 common convolution, a first batch normalization layer, an average pooling layer, and a second batch normalization layer; the fourth branch is sequentially processed by a k×k common convolution and a batch normalization layer; the fifth branch is sequentially processed by a 1×k common convolution and a batch normalization layer; the sixth branch is sequentially processed by a k×1 common convolution and a batch normalization layer; the features processed by the six branches are convolved and summed, and the output features are obtained after passing through the SiLU activation function; The processing process of the feature fusion network is as follows: the output feature F3 of the PSA module is upsampled by the first upsampling module and then fused with the output feature F2 of the third CWDB module through the first fusion module to obtain the fused feature C1, which is input to the fifth CWDB module for processing; After the output feature C2 of the fifth CWDB module is processed by the upsampling module, it is fused with the output feature F1 of the second CWDB module through the second fusion module to obtain the fused feature C3; the fused feature C3 is processed by the sixth CWDB module, and the output feature C4 of the sixth CWDB module is processed by the fourth DRFD module and then fused with the output feature C2 of the fifth CWDB module through the third fusion module to obtain the fused feature C5. The fused feature C5 is input into the seventh CWDB module for processing. The output feature C6 of the seventh CWDB module is processed by the fifth DRFD module and then fused with F3 through the fourth fusion module to obtain the fused feature C7. The fused feature C7 is input into the C2FCIB module for processing. The output feature C4 of the sixth CWDB module, the output feature C6 of the seventh CWDB module, and the output feature C8 of the C2FCIB module are used as the inputs of the three detection modules of the detection network.
2. The method for identifying birds hitting power lines according to claim 1, wherein The operation mode of the PSA module is as follows: the module is divided into seven parts, namely the first ordinary convolution, the multi-head self-attention module, the first feature splicing module, the feed-forward network, the second feature splicing module, the fusion module, and the second ordinary convolution. The output features of the first ordinary convolution are respectively input into the multi-head self-attention module, the first feature splicing module, and the fusion module. The spliced features of the first feature splicing module are respectively input into the feed-forward network and the second feature splicing module. After the feature fusion by the fusion module, it is input into the second ordinary convolution for convolution processing to obtain the final output of the PSA module.
3. The method for identifying birds hitting power grids according to claim 1, wherein The processing process of the C2FCIB module is as follows: the image features are input into a CBS module for convolution processing, and then after being processed by the separation module, they are successively processed by several CIB modules. The CIB module is composed of three depthwise separable convolutions and two CBS modules connected in series in sequence. The connection method is the first depthwise separable convolution, the first CBS module, the second depthwise separable convolution, the second CBS module, and the third depthwise separable convolution. Finally, the features of the separation module and several CIB modules are fused through the fusion module, and the fused features are input into another CBS module for processing to obtain the output of the C2FCIB module.
4. The method for identifying birds hitting power lines according to claim 1, characterized in that The shallow feature extraction module is the SRFD module. The SRFD module includes three stages. The first stage uses ordinary convolution; the features output in the first stage enter the second stage. The upper branch of the second stage consists of grouped convolution and depthwise separable convolution, and the lower branch of the second stage is pixel cutting downsampling. After the features of the upper branch and the lower branch are spliced and convolved, they enter the third stage; in the upper branch of the third stage, first grouped convolution is performed, and then depthwise separable convolution and max-pooling downsampling are respectively performed through two branches. The features output by the depthwise separable convolution and the max-pooling downsampling are spliced and convolved with the output of the lower branch of the third stage, and finally the output features of the SRFD module are obtained.
5. The method for identifying birds hitting power lines according to claim 1, wherein The detection module consists of two branches. The operation mode of the first branch is as follows: after the input features are convolved by the CBS module, they are successively processed by the MultiSEAM module and ordinary convolution to obtain the prediction loss. The operation mode of the second branch is as follows: after the input features are processed by two consecutive CBS modules, they are successively processed by the MultiSEAM module and ordinary convolution to obtain the classification loss. Among them, the operation mode of the MultiSEAM module is as follows: the original input features are respectively input into three CSM modules for processing. The output features obtained by the three CSM modules are convolved and summed with the original input features. After the summation, the features are respectively subjected to average pooling and full connection, and the obtained output is multiplied by the original input features to obtain the final output features.
6. A power grid collision line bird recognition system based on multi-scale feature enhancement, for implementing the power grid collision line bird recognition method according to claim 1, characterized in that, It includes an image acquisition device and a power grid wire-crossing bird recognition device. The image acquisition device is used to acquire bird images within a certain distance range of the transmission line. The power grid wire-crossing bird recognition device is built-in with an RBOD rare bird target detection model for bird recognition.
Citation Information
Patent Citations
Power line bird identification method based on lightweight target detection model
CN117237871A
Single-stage target detection method for traffic sign detection
CN118015581A