Remote Sensing Target Detection Method for Electric Power Transmission Towers Based on Large-Core Selection Feature Fusion Network

By improving the backbone and neck network of the YOLOv5 model, combined with multi-scale feature fusion and improved loss function, the problem of insufficient accuracy of power tower detection in satellite remote sensing images is solved, and high-precision detection and classification of power towers is achieved.

CN118212546BActive Publication Date: 2025-07-29NORTH CHINA ELECTRIC POWER UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410208305.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-26
Publication Date
2025-07-29
Estimated Expiration
2044-02-26

AI Technical Summary

Technical Problem

In satellite remote sensing images, the detection of power towers has problems such as insufficient detection accuracy, small target area, inconsistent scale and complex background, especially insufficient detection accuracy for distribution towers.

Method used

The large-core selection feature fusion network is adopted, and the receptive field is expanded by improving the backbone and neck networks of the YOLOv5 model, and the multi-scale feature alignment fusion structure is used, and the MPDIoU loss function and sliding weighted loss are combined to improve the detection accuracy of the model for the power tower.

Benefits of technology

It effectively improves the detection accuracy of power pole towers, especially for the tiny pole towers of power distribution towers, achieving accurate positioning and classification in complex contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118212546B_ABST
    Figure CN118212546B_ABST
Patent Text Reader

Abstract

The present invention discloses a remote sensing target detection method for power transmission towers based on a large kernel selection feature fusion network, belonging to the field of power computer vision. The method includes the steps of: selecting YOLOv5 as the basic model, designing a large kernel spatial selection attention fusion module to improve the backbone network, expanding the receptive field of the model, and accurately positioning the location of the power transmission tower. Designing a multi-scale feature alignment and fusion structure to improve the neck network, solving the problem of large and inconsistent scale gaps between transmission towers and distribution towers, improving the detection accuracy of such small towers as distribution towers, and realizing multi-scale feature fusion of power transmission towers under complex backgrounds. Training the model, introducing MPDIoU to improve CIoU, designing a sliding weighted loss to make the model pay more attention to negative samples during the training process, using the optimal model obtained from training to detect and identify power transmission towers, and evaluating the model effect. The present invention effectively improves the detection accuracy of power transmission towers in satellite remote sensing images, and can provide important technical support for the intelligent inspection of power lines based on satellite remote sensing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power vision target detection, and particularly to a method for remotely sensing target detection of power transmission towers based on a large kernel selection feature fusion network. Background Art

[0002] With the continuous development and application of artificial intelligence technology, the inspection of power lines has gradually shifted from digital to intelligent. As one of the important power infrastructure, power transmission towers are mainly used to support power lines to ensure the safe operation of power lines, and at the same time ensure a certain distance between the power lines and the ground. Compared with the inspection by unmanned aerial vehicles, the intelligent inspection of power lines based on satellite remote sensing can achieve large-scale and business-oriented inspection of power corridors, which can greatly improve the inspection efficiency and pertinence, and has gradually become an important research direction. Using computer vision and image processing technologies to obtain the position information of power transmission towers from satellite images can provide services for power companies to inspect the distribution of electric towers, and better assist inspectors to carry out their work, providing assistance for planning regional power grids and monitoring the vegetation erosion of power corridors.

[0003] Some existing studies have used computer vision technology to identify power transmission towers in satellite inspection images. For example, in the patent "Method for Remotely Sensing Target Detection of Transmission Towers Based on Semi-Supervised Learning and Deformable Convolution", by applying semi-supervised learning and deformable convolution to the remotely sensing target detection of transmission towers, an ENGD-BiFPN (EfficientNet GroupedDeformable BiFPN) detection model is constructed to more accurately detect the transmission tower targets in satellite remote sensing images and improve the survey efficiency of transmission towers. However, there are still the following two problems in using deep learning target detection methods to detect power transmission towers in satellite remote sensing images:

[0004] 1. In satellite remote sensing images, the pixel area occupied by power transmission towers is less compared to the background, making it difficult to extract effective features. Moreover, the sizes and scales of the two types of targets, namely distribution towers and transmission towers, are different and unevenly distributed, resulting in difficulty in improving the detection accuracy.

[0005] 2. The resolution of satellite remote sensing images is inherently insufficient for target detection. The deepening of the number of layers in the CNN model causes the resolution of the images to continuously decrease, and many detailed features become more blurred after multiple convolutional operations. The neglect of low-level feature information leads to insufficient detection accuracy for small-scale towers such as distribution towers.

[0006] 3. The background of satellite remote sensing images is relatively complex, and many power transmission towers have similar features and textures to the surrounding ground objects, resulting in difficulty in improving the detection accuracy. Summary of the Invention

[0007] To address the deficiencies of existing satellite remote sensing image power pole detection technologies, the motivation of the present invention is to propose an object detection framework that can accurately locate the positions of power poles in satellite remote sensing images and distinguish whether they belong to transmission towers or distribution towers, providing important technical support for intelligent inspection of power lines based on satellite remote sensing. The present invention proposes a large kernel selection feature fusion network for power pole detection in high-resolution satellite remote sensing images, effectively improving the detection accuracy of power poles in satellite remote sensing images.

[0008] This application proposes a power pole remote sensing object detection method based on a large kernel selection feature fusion network, including the following steps:

[0009] S1: Construct a satellite image power pole dataset. Based on the collected satellite remote sensing images of transmission and distribution infrastructure, obtain the labeled images of power poles covering different backgrounds and different population density regions, where the power poles include transmission towers and distribution towers.

[0010] S2: Weighing accuracy and speed, select YOLOv5 as the basic model, design a large kernel spatial selection attention fusion module to improve the backbone network, expand the model's receptive field, and accurately locate the positions of power poles.

[0011] S3: Design a multi-scale feature alignment and fusion structure to improve the neck network, effectively utilize low-level semantic information, solve the problem of large and inconsistent scale gaps between transmission towers and distribution towers, improve the detection accuracy of small power poles such as distribution towers, and achieve multi-scale feature fusion of power poles in complex backgrounds.

[0012] S4: Train the model, introduce MPDIoU to improve CIoU, design a sliding weighted loss to make the model pay more attention to negative samples during the training process, use the optimal model obtained from training to detect and identify power poles, and evaluate the model's performance.

[0013] Compared with the existing technology, the beneficial effects of the present invention are as follows: By adding a large kernel spatial selection attention mechanism to the backbone network of the YOLOv5 benchmark model, the receptive field of the model is expanded, the ability of the model to extract context is enhanced, and the attention feature fusion module is used to improve its residual connection to prevent the problem of gradient disappearance caused by the deepening of the model network layers, realizing the accurate positioning of power poles in complex backgrounds; A multi-scale feature fusion neck network is constructed, which can prevent the loss of small target feature information caused by the deepening of the model layers, make more effective use of low-level semantic information containing more detailed features, realize the multi-scale fusion of power poles, and improve the detection accuracy of small power poles such as distribution towers; Use MPDIOU to improve CIoU to better distinguish between two types of power poles, namely transmission towers and distribution towers; Compared with the existing satellite remote sensing power pole detection methods, the model trained in this paper effectively improves the power pole detection accuracy and can automatically detect and classify transmission and distribution towers. Brief Description of the Drawings

[0014] The accompanying drawings that form a part of the present invention are used to provide a further understanding of the present invention.

[0015] Figure 1 It is a flowchart of a remote sensing target detection method for power transmission towers based on a large kernel selection feature fusion network proposed by the present invention;

[0016] Figure 2 It is a schematic diagram of the large kernel selection feature fusion network LSKF-YOLO in an embodiment of the present invention;

[0017] Figure 3 It is a schematic diagram of data processing and annotation in an embodiment of the present invention;

[0018] Figure 4 It is a schematic diagram of the large kernel selection attention feature fusion module LSKM in an embodiment of the present invention;

[0019] Figure 5 It is a schematic diagram of the multi-scale feature fusion structure MFAF in an embodiment of the present invention;

[0020] Figure 6 It is a schematic diagram of the structure of the SPD-Conv module in an embodiment of the present invention;

[0021] Figure 7 It is the schematic diagram of MPDIoU in an embodiment of the present invention;

[0022] Figure 8 It is the effect diagram of satellite remote sensing power transmission tower detection in an embodiment of the present invention. Detailed Embodiment

[0023] Next, the technical content, implementation purpose and effects in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0024] This embodiment provides a remote sensing target detection method for power transmission towers based on a large kernel selection feature fusion network. Please refer to Figure 1 、 Figure 2 , this method includes:

[0025] S1: Construct a satellite image power transmission tower data set. According to the collected satellite remote sensing images of power transmission and distribution infrastructure, obtain the labeled images of power transmission towers covering different backgrounds and different population density areas, where the power transmission towers include transmission towers and distribution towers;

[0026] S2: Weighing accuracy and speed, select YOLOv5 as the basic model, design a large kernel spatial selection attention fusion module to improve the backbone network, expand the model receptive field, and accurately locate the position of the power transmission tower;

[0027] S3: Design a multi-scale feature alignment and fusion structure to improve the neck network, effectively utilize low-level semantic information, solve the problem of large and inconsistent scale gaps between transmission towers and distribution towers, improve the detection accuracy of such small-scale poles and towers as distribution towers, and achieve multi-scale feature fusion of power poles and towers under complex backgrounds;

[0028] S4: Train the model, introduce MPDIoU to improve CIoU, design a sliding weighted loss to make the model pay more attention to negative samples during training, use the optimal model obtained from training to detect and identify power poles and towers, and evaluate the model effect;

[0029] Further, step S1 includes:

[0030] S11: Crop the collected satellite images of transmission and distribution infrastructure into images of size 512×512 from top to bottom and from left to right, clean the obtained images, and select the images with clear images, more types and numbers of power poles and towers to summarize into a picture subset.

[0031] S12: Use LabelImg to re-label the picture subset. The labeling categories are divided into two categories: transmission towers and distribution towers. For such small-scale poles and towers as distribution towers, label their shadows and themselves as a single unit to obtain a labeled satellite remote sensing power pole and tower dataset.

[0032] Specifically, please refer to Figure 3 :

[0033] Deep learning object detection networks are limited by the size of the images in the dataset. In the original dataset of this example, the resolution of the images is between 3800 and 12000 pixels, and the aspect ratio of some images is too large to be directly trained. To better meet the needs of model training, the images in the original dataset are all cropped into images of size 512×512 from top to bottom and from left to right. In addition, since the annotations in the original dataset are polygon annotations, which are not applicable to the object detection model, we use labelImg to re-label the images. The labeling categories are divided into two categories: transmission towers and distribution towers. Since it is difficult to observe distribution towers in satellite images, label their shadows and themselves as a single unit.

[0034] Further, step S2 includes:

[0035] S21: Select YOLOv5 with a small model size and fast detection speed as the benchmark model. The backbone network CSPDarknet53 of the model consists of Conv, C3, SPPF and bottleneck modules. Add a large kernel spatial selection attention mechanism module LSKM (Large Selective Kernel Mechanism) before the SPPF layer to enhance the extraction ability of the model backbone network.

[0036] The Large Kernel Spatial Selection Attention Mechanism (LSKM) network structure in step S21 includes: a Large Kernel Spatial Selection sub-block and a Multi-Layer Perceptron (MLP) sub-block. The core of the Large Kernel Spatial Selection sub-block is the LSK module, which consists of a large kernel convolution sequence and a spatial kernel selection mechanism. The large kernel convolution sequence is composed of a depth convolution sequence with a large growth kernel and an increasing dilation rate through explicit decomposition. The spatial kernel selection mechanism performs spatial selection on the feature map from large convolution kernels of different scales.

[0037] Specifically, please refer to Figure 4 :

[0038] The decomposed kernel of the large kernel convolution sequence is denoted as D i (D i = X). Assuming there are N decomposed kernels, dw i represents a depth convolution with kernel K i and dilation rate d i . After each kernel is decomposed, it is processed by a 1×1 convolution layer Cl 1×1 . The decomposition process is shown in the following formula:

[0039]

[0040] The spatial selection mechanism is used to improve the network's ability to focus on the most relevant spatial context regions to enhance target localization. Its steps include: First, the features obtained from different kernels are concatenated with different receptive field ranges to obtain a feature map Then, apply channel-based average pooling P and max pooling P avg and P max to obtain a spatial attention map Next, for each , they are concatenated and Cl 2→N is used to change the 2 channels to N channels. Cl 2→N is a 1×1 convolution layer, and then a separate spatial selection mask for each decomposed large kernel is obtained through the Sigmoid activation function Next, the features in the decomposed kernels are weighted accordingly and fused through a convolution layer to obtain the attention feature S. The final output Y is the fusion of the input feature map and the attention feature. The above process can be expressed as follows:

[0041]

[0042]

[0043]

[0044] Y = X · S (5)

[0045] S22: Replace the residual connection between each sub - layer of LSKM with the attention feature fusion module AFF (Attention Feature Fusion) to avoid the problems of gradient disappearance and weight matrix degradation caused by the increase in the number of network layers, enhance the ability to capture different local information, and improve the performance of the backbone network in extracting targets under complex backgrounds.

[0046] Furthermore, step S3 includes:

[0047] S31: In the neck network of the baseline model YOLOv5, use PANet. To better fuse low - level semantic information, replace the first two modules of the neck network with the multi - scale feature alignment and fusion structure MFAF (Muti - scale Feature Alignment and Fusion).

[0048] The MFAF in step S31 includes: the feature alignment module FAM (Feature Alignment Model), the information fusion module IFM (Information Feature Fusion), and the lightweight adjacent layer fusion module LAF_Injection Module (Lightweight Adjacent Layer Fusion Model).

[0049] S32: Use SPD - conv instead of the original strided convolutional layer in the last two neck network modules to prevent the feature map from blurring due to the increase in network depth, which affects the detection effect of the detection head, and improve the detection ability for small - target poles such as distribution towers.

[0050] Specifically, please refer to Figure 2 、 Figure 5 :

[0051] In the feature alignment module (FAM), select Figure 2 the output features P1, P2, P3, P4 of the backbone network in Figure 2 for alignment and fusion to obtain high - resolution features that retain small - target information. Use average pooling (AvgPool) operation to downsample the input features and achieve a unified size. Select ali .

[0052] X ali = FAM([P1, P2, P3, P4]) (6)

[0053] X fus= RepBlock(X ali ) (7)

[0054] X inj_I , X inj_II = Split(X fus ) (8)

[0055] Then, attention operations are adopted in the LAF_Injection_Module to fuse the information, which includes two parts: local feature x_local and global information x_global. The local feature of LAF_Injection_Module I comes from the P2, P3, and IFM information fused by LAF, and the local feature of LAF_Injection_Module II comes from the IFM, P3, and P4 information fused by LAF. The global feature is the output information of IFM. x_global is calculated using two different convolutional layers to obtain the features of two branches, and x_local is calculated only through one convolutional layer. Then, the features of these three outputs are fused through attention calculation. Due to the difference between the local feature and the global feature, average pooling or bilinear interpolation is used to scale the output features according to the size of the local information to keep them aligned. After the fusion is completed, RepBlock is added to further extract the fused information.

[0056] X global_out1 = resize(sigmoid(Conv(X_global)) (9)

[0057] X global_out2 = resize(Conv(X_global)) (10)

[0058] X attn = Conv(X_local)*X global_out1 + X global_out2 (11)

[0059] X out = RepBlock(X attn ) (12)

[0060] The adopted SPD-Conv consists of a space-to-depth convolutional layer SPD and a non-strided convolution. Please refer to Figure 6 , specifically including:

[0061] For the input feature X with a scale of S×S and a channel dimension of C, the scaling factor is set to 2 to obtain four sub-feature maps with their scales halved to S / 2 and the channel dimension remaining unchanged. Then, these feature sub-maps are concatenated in the channel dimension, so that the channel dimension becomes 4C. Finally, an output feature map is obtained through a non-strided convolution with a stride of 1.

[0062] Further, step S4 includes:

[0063] S41: Divide the data set into a training set and a validation set according to a ratio of 8:2, where the validation set is also the test set, and use the 5-k cross-validation method to train the data set. The 5-k cross-validation method alternately uses the data set as the training set and the validation set, which can effectively solve the problem of class imbalance of transmission towers and distribution towers;

[0064] S42: In the training loss, Focal_Loss is still used to calculate the object and classification loss, and MPDIoU is introduced to improve CIoU as the localization loss. MPDIoU is a new boundary box similarity comparison metric based on the minimum point distance. To solve the problem that when the predicted box and the true box have the same aspect ratio but completely different width and height values, the existing boundary box regression loss function cannot be optimized. Inspired by the geometric characteristics of the boundary box, MPDIoU forces each boundary box predicted by the model to approach its true box by minimizing the loss function during the training phase. And the four point coordinates of the boundary box are used to represent all factors of the existing boundary box regression loss function. MPDIoU can briefly represent all factors considered in the existing loss function, covering overlapping and non-overlapping regions, center point distance, width and height deviation, while simplifying the calculation process.

[0065] Specifically, please refer to Figure 7 , and its calculation process is as follows:

[0066] [[ID=********]]

[0067] [[ID=********]] [[ID=********]] [[ID=********]]

[0068] [[ID=********]] [[ID=********]] [[ID=********]]

[0069] L MPDIoU = 1 - MPDIoU (16)

[0070] where w and h represent the width and height of the input image, represents the coordinates of the upper left and lower right corners of the true box, represents the coordinates of the upper left and lower right corners of the predicted box, d1 2 represents the upper left distance, d2 2 lower right distance, L MPDIoU represents the MPDIoU_Loss calculated by MPDIoU. Please refer to Figure 7 .

[0071] Thus, all factors of the existing boundary box regression loss function can be determined by four point coordinates. The conversion formula is as follows:

[0072]

[0073]

[0074]

[0075] where |C| represents the area of the smallest enclosing rectangle covering A and B, representing the coordinates of the centers of the ground truth box and the predicted box respectively. w A , h A represent the width and height of the ground truth box, and w B , h B represent the width and height of the predicted box;

[0076] S43: To make the model pay more attention to negative samples during training, we add a sliding weighted loss in the experiment. The average value of MPDIoU of all obtained bounding boxes is used as the threshold. Those less than the threshold are used as negative samples, and those greater than the threshold are used as positive samples. Then, a higher weight is assigned to the samples near the threshold through the weighted function Slide, and the Slide function is expressed as follows

[0077]

[0078] S44: According to the optimal model obtained from training, the power transmission towers are detected and recognized, and the model performance is evaluated. The evaluation metrics include precision, recall, and mean average precision (mAP). The calculation process is as follows:

[0079]

[0080] [[ID=3e]]

[0081]

[0082]

[0083] where TP represents the number of detected real power transmission towers, FP represents the number of wrongly detected power transmission towers, and FN represents the number of undetected power transmission towers.

[0084] By comparing Faster R-CNN, SSD, RetinaNet, and other models in the YOLO series (YOLOv4, YOLOv5s), the advantages of the LSKF-YOLO model of the present invention can be illustrated. Table 1 shows the comparison between LSKF-YOLO and the above models, and it can be seen that LSKF-YOLO performs well in terms of the mAP score and the accuracy metric.

[0085] Table 1 Comparison of Model Results

[0086]

[0087] The detection effect of the method of the present invention is as Figure 8 shown. Based on the YOLOv5 model and the idea of large convolutional kernel attention, the present invention effectively fuses low-level feature information by using a feature alignment and fusion module to achieve multi-scale power tower detection. First, a large kernel spatial selection attention fusion structure LSKM is added to the backbone feature extraction layer to expand the network receptive field and enhance the feature extraction ability of the backbone network. Second, a multi-scale feature alignment and fusion MFAF structure is introduced into the feature fusion layer of the neck network to enhance the feature representation in the network and effectively utilize low-level semantic information. Finally, SPD-Conv is used to replace the strided convolution of the last two layers of the neck network, enhancing the detection ability of small targets such as distribution towers in the image. Experimental results on the constructed satellite remote sensing power tower dataset show that the method has good detection performance for power tower targets in satellite remote sensing images with complex backgrounds.

[0088] The specific examples described in the present invention are only illustrative of the core idea and spirit of the present invention. For those of ordinary skill in the art, there will be changes in the specific implementation manners and application scopes according to the idea of the present invention. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A remote sensing target detection method for power transmission towers based on a large-core selection feature fusion network, characterized in that It includes the following steps: S1: Construct a satellite image power pole dataset. Based on the collected satellite remote sensing images of transmission and distribution infrastructure, obtain the labeled images of power poles covering different backgrounds and regions with different population densities. Among them, the power poles include transmission towers and distribution towers; S2: Weighing accuracy and speed, select YOLOv5 as the basic model, design a large kernel spatial selection attention fusion module to improve the backbone network, expand the model's receptive field, and accurately locate the position of power poles; S3: Design a multi-scale feature alignment and fusion structure to improve the neck network, effectively utilize low-level semantic information, solve the problem of large and inconsistent scale gaps between transmission towers and distribution towers, improve the detection accuracy of such small power poles as distribution towers, and achieve multi-scale feature fusion of power poles under complex backgrounds; S4: Train the model, introduce MPDIoU to improve CIoU, design a sliding weighted loss to make the model pay more attention to negative samples during training, use the optimal model obtained from training to detect and identify power poles, and evaluate the model's performance; Among them, in step S2, YOLOv5 is selected as the basic model, and a large kernel spatial selection attention fusion module is designed to improve the backbone network, which specifically includes: S21: Select YOLOv5 with a small model size and fast detection speed as the benchmark model. The backbone network CSPDarknet53 of the model is composed of Conv, C3, SPPF, and bottleneck modules. Add a large kernel spatial selection attention mechanism module LSKM (Large Selective Kernel Mechanism) before the SPPF layer to enhance the extraction ability of the model's backbone network; The network structure of the large kernel spatial selection attention mechanism LSKM in step S21 includes: a large kernel spatial selection sub-block (Large Kernel Spatial Selection) and a feedforward neural network sub-block (MLP). The core of the large kernel spatial selection sub-block is the LSK module, which is composed of a large kernel convolution sequence and a spatial kernel selection mechanism. The large kernel convolution sequence is composed of a depth convolution sequence with a large growth kernel and an increasing dilation rate through explicit decomposition. The spatial kernel selection mechanism performs spatial selection on the feature map from large convolution kernels of different scales; S22: Use the attention feature fusion module AFF (Attention Feature Fusion) to replace the residual links between each sub-layer of LSKM, avoid the problems of gradient disappearance and weight matrix degradation caused by the increase in the number of network layers, enhance the ability to capture different local information, and improve the performance of the backbone network in extracting targets under complex backgrounds; Among them, in step S3, a multi-scale feature alignment and fusion structure SPD-MFAF is designed to improve the neck network, which specifically includes: S31: The neck network of the benchmark model YOLOv5 uses PANet. In order to better fuse low-level semantic information, use a multi-scale feature alignment and fusion structure MFAF (Muti-scale Feature Alignment and Fusion) to replace the first two modules of the neck network; The MFAF in step S31 includes: a Feature Alignment Module (FAM), an Information Feature Fusion Module (IFM), and a Lightweight Adjacent Layer Fusion Model (LAF_Injection Module). S32: Therefore, SPD-conv is used to replace the original strided convolutional layer in the latter two neck network modules to prevent the feature map from being blurred due to the increase in network depth, which may affect the detection effect of the detection head and improve the detection ability for small power towers such as distribution towers.

2. The method for remotely sensing target detection of power transmission towers based on a large-core selection feature fusion network according to claim 1, wherein The method for constructing the satellite remote sensing power tower dataset in step S1 specifically includes: S11: The collected satellite images of the power transmission and distribution infrastructure are cropped into images of size 512×512 from top to bottom and from left to right. The obtained images are cleaned, and the images with clear images, more types and numbers of power towers are selected and summarized to obtain a subset of images. S12: Use LabelImg to re-label the subset of images. The labeled categories are divided into two types: transmission towers and distribution towers. For relatively small power towers such as distribution towers, their shadows and themselves are labeled as a single unit to obtain the labeled satellite remote sensing power tower dataset.

3. The method for remote sensing target detection of power transmission towers based on the large-core selection feature fusion network according to claim 1, wherein, In step S4, the model is trained, and the trained optimal model is used to detect and identify power towers, and the model effect is evaluated, which specifically includes: S41: The dataset is divided into a training set and a validation set (the validation set is also the test set) in a ratio of 8:

2. The 5-k cross-validation method is used to train the dataset to solve the problem of class imbalance between transmission towers and distribution towers. S42: In the training loss, Focal_Loss is still used to calculate the object and classification losses, and MPDIoU is introduced to improve CIoU as the localization loss. MPDIoU is a new boundary box similarity comparison metric based on the minimum point distance. This metric includes all factors considered in the existing loss functions, covering overlapping and non-overlapping regions, the distance between the center points, and the deviations of width and height, while simplifying the calculation process. S43: To make the model pay more attention to negative samples during training, a sliding weighted loss is added in the experiment. The average value of MPDIoU of all obtained bounding boxes is used as the threshold. Samples smaller than the threshold are used as negative samples, and samples larger than the threshold are used as positive samples. Then, a higher weight is assigned to the samples near the threshold through the weighting function Slide. S44: According to the trained optimal model, power towers are detected and identified, and the model performance is evaluated. The evaluation indicators include precision, recall, and mean average precision (mAP).

Citation Information

Patent Citations

  • Remote sensing image transmission tower detection method based on feature enhanced convolutional network

    CN114022764A

  • Target tracking method, system and device for multi-view image clustering and medium

    CN117292162A