Aerial photography insulator defect detection method based on multi-scale fusion and cascade detection

By adopting an improved ReDet model based on multi-scale fusion and cascade detection in drone inspection, the problem of small and medium-sized target and multi-object detection insulators and their defects is solved, which significantly improves detection accuracy and robustness, and is suitable for drone inspection in complex contexts.

CN120219295APending Publication Date: 2025-06-27SOUTHEAST UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510247560.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The insulators and defect detection in existing drone inspection images are difficult to solve the problems of small-target and multi-target detection, and the detection accuracy and robustness are insufficient in complex backgrounds.

Method used

Aerial insulator defect detection method based on multi-scale fusion and cascade detection is adopted to improve detection accuracy and robustness through an improved ReDet model, including a feature extraction module, an improved feature pyramid module (ReSiRFP) and a cascading rotation-invariant RoI detector (CRiRoI Head).

Benefits of technology

In the complex context, the detection accuracy and robustness of rotating small targets is significantly improved, and the missed detection and false detection rates are reduced, making it suitable for deployment on resource-constrained drone platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219295A_ABST
    Figure CN120219295A_ABST
Patent Text Reader

Abstract

The invention discloses an aerial insulator defect detection method based on multi-scale fusion and cascade detection, which comprises the following steps of: firstly, preprocessing an aerial insulator image data set, inputting the aerial insulator image data set into an improved ReDet model, extracting rotation isovariant features of an input image through a backbone network feature extraction module in the model, and extracting the rotation isovariant features of the input image through a backbone network feature extraction module; generating multi-scale feature information by using a rotation equivariant recursive feature pyramid constructed based on a multi-scale fusion strategy; and a candidate box is generated through a region suggestion network, the candidate box is input into a cascaded rotation invariant RoI detector, target classification and boundary regression are refined stage by stage, and finally an image containing target category and position information is output, so that high-precision detection of the rotating small target under a complex background is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information security, and mainly relates to an aerial insulator defect detection method based on multi-scale fusion and cascaded detection. Background Art

[0002] The transmission line is an important part of the power grid, and insulators undertake the functions of wire support and electrical insulation. During long-term operation, insulators are prone to defects due to mechanical tension and high voltage, which may cause power outages. Therefore, the efficient inspection and accurate detection of insulators are crucial for the power system.

[0003] In recent years, due to its advantages of high efficiency, safety, and low cost, the unmanned aerial vehicle (UAV) has become an important tool for the inspection of transmission lines. Equipped with a high-definition camera and sensors, the UAV can obtain high-precision images in complex environments and improve the inspection efficiency. Object detection methods based on deep learning have been widely used in insulator defect detection, and the mainstream algorithms include SSD, YOLO series, and FasterR-CNN, etc. However, traditional horizontal bounding box detection methods face problems such as the lack of target direction information and excessive background information in the detection box in the UAV inspection scenario, which affect the detection accuracy. The method based on rotated bounding boxes has made some progress in remote sensing images, but still faces challenges such as insufficient detail extraction and target occlusion in UAV inspections, which limit the accuracy and robustness of detection.

[0004] Therefore, it is of great research and practical significance to develop a high-precision insulator and defect detection algorithm suitable for UAV inspections to solve the problems of small target and multi-target detection. Summary of the Invention

[0005] In view of the problems that the insulators and their defects in the existing UAV inspection images have large target scale variations, arbitrary rotation angles, complex backgrounds, etc., the present invention proposes an aerial insulator defect detection method based on multi-scale fusion and

[0006] cascaded detection. First, the aerial insulator image dataset is preprocessed and input into an improved ReDet model. In this model, the rotated equivariant features of the input image are first extracted by the backbone network feature extraction module

[0007] and multi-scale feature information is generated by using a rotated equivariant recursive feature pyramid constructed based on a multi-scale fusion strategy; then candidate boxes are generated by the region proposal network, and the candidate boxes are input into the cascaded rotation-invariant RoI detector for stage-by-stage refinement of target classification and boundary regression. Finally,

[0008] an image containing target category and position information is output, realizing high-precision detection of rotated small targets in complex backgrounds.

[0009] To achieve the above object, the technical solution adopted by the present invention is: an aerial insulator defect detection method based on multi-scale fusion and cascaded detection, including the following steps:

[0010] S1. Data preprocessing: Preprocess the data in the UAV aerial insulator dataset, and

[0011] annotate and partition the preprocessed dataset, and wait to enter the improved ReDet model; the improved ReDet model at least includes a feature extraction module, an improved feature pyramid module (ReSiRFP), a region proposal network, and an improved RoI detector (CRiRoI Head); load the pre-trained ReDet model as the basic network, and improve it, replace the original feature pyramid module with the ReSiRFP module, and replace the RoI detector with the CRiRoI Head to complete the initialization settings of the model, and provide a basic framework for subsequent training

[0012] and detection;

[0013] S2. Feature extraction: Send the data preprocessed in step S1 into the improved ReDet model, and perform feature extraction in the feature extraction module to generate a feature map with rotational equivariance;

[0014] S3. Obtaining multi-scale feature maps: Input the feature map with rotational equivariance generated in step S2 into the improved feature pyramid module to obtain multi-scale feature maps; the improved feature pyramid module at least includes

[0015] a rotation-equivariant recursive feature pyramid network, a rotation-equivariant atrous spatial pyramid pooling module, and a Fusion module;

[0016] S4. Processing of the region proposal network: Send the multi-scale feature maps obtained in step S3 into the region proposal network for processing. The region proposal network at least includes a shared convolutional layer, a classification branch, and a bounding box regression branch. The classification branch and the bounding box regression branch are two parallel fully connected branches; use the sliding window mechanism to

[0017] slide on the feature map to generate a series of candidate regions, and screen out potential target region candidate boxes through classification and bounding box regression operations;

[0018] S5. Processing of the improved RoI detector: Input the target region candidate boxes obtained in step S4 into the improved RoI detector, perform spatial and orientation feature alignment, and then predict the target category through the classification branch and adjust the position of the candidate box through the bounding box regression branch to generate a detection image containing target category and position information.

[0019]

[0020] As an improvement of the present invention, the data preprocessing in step S1 includes at least random rotation, cropping, scaling, and brightness adjustment to enhance data diversity and the generalization ability of the model; accurately label insulators and their defects in the dataset, and divide the labeled data into a training set and a test set according to a ratio of 9:1 for model training and performance evaluation respectively.

[0021] After the input image is preprocessed in step S2, it is sent into the backbone network ReResNet of the improved model. ReResNet performs preliminary convolutional feature extraction on the input image, extracts features at a fixed number of rotation angles through rotation-equivariant convolution, and combines the weight sharing mechanism to ensure the consistency of features under rotation changes. Then, the features are further enhanced and fused in the residual block to generate a feature map with rotation equivariance; the generated rotation-equivariant feature map has a size of (K, N, H, W), where K is the number of channels, N is the number of direction channels, and H and W are the spatial dimensions of the feature map.

[0022] As another improvement of the present invention, the acquisition of the multi-scale feature map in step S3 specifically includes the following steps:

[0023] S31: The input feature map enters the core part of the improved feature pyramid (ReSiRFP), the rotation-equivariant recursive feature pyramid network, and generates a preliminary multi-scale feature map through bottom-up and top-down feature transfer;

[0024] S32: The preliminary multi-scale feature map passes through the rotation-equivariant atrous spatial pyramid pooling module (ReSiASPP) in the feedback link of ReSiRFP, and realizes rotation-equivariant multi-scale feature extraction through atrous convolutions with different dilation rates, and dynamically adjusts the feature response through the self-gated activation function ennSwish;

[0025] S33: After the multi-scale feature map passes through the bottom-up and top-down transfer link again, it passes through the Fusion module in the output link of ReSiRFP, and dynamically adjusts the contribution degrees of features at each scale through the adaptive weight mechanism to output the final multi-scale feature map; the specific method of dynamically adjusting the contribution degrees of features at each scale by using the adaptive weight mechanism is as follows:

[0026]

[0027] As another improvement of the present invention, the acquisition of the multi-scale feature map in step S3 specifically includes the following steps:

[0028] S31: The input feature map enters the core part of the improved feature pyramid (ReSiRFP), the rotation-equivariant recursive feature pyramid network, and generates a preliminary multi-scale feature map through bottom-up and top-down feature transfer;

[0029] S32: The preliminary multi-scale feature map passes through the rotation-equivariant atrous spatial pyramid pooling module (ReSiASPP) in the feedback link of ReSiRFP, and realizes rotation-equivariant multi-scale feature extraction through atrous convolutions with different dilation rates, and dynamically adjusts the feature response through the self-gated activation function ennSwish;

[0030] S33: The multi-scale feature map passes through the rotation-equivariant atrous spatial pyramid pooling module (ReSiASPP) in the feedback link of ReSiRFP, and realizes rotation-equivariant multi-scale feature extraction through atrous convolutions with different dilation rates, and dynamically adjusts the feature response through the self-gated activation function ennSwish;

[0031] S33: After the multi-scale feature map passes through the bottom-up and top-down transfer link again, it passes through the Fusion module in the output link of ReSiRFP, and dynamically adjusts the contribution degrees of features at each scale through the adaptive weight mechanism to output the final multi-scale feature map; the specific method of dynamically adjusting the contribution degrees of features at each scale by using the adaptive weight mechanism is as follows:

[0032]

[0033] Among them, F rot-equiv is a rotation-equivariant feature map, ennConv is a rotation-equivariant convolution operation, and F multi-scale is a multi-scale feature map obtained through a feature pyramid network, W is the weight map generated by F rot-equiv σ is the sigmoid function, and ennAdaptiveAvgPool is a rotation-equivariant average pooling operation. F fused is the feature map output after weighted fusion, and F curr is the feature at the current step, and F prev is the feature at the previous step. ⊙ represents element-wise multiplication.

[0034] As another improvement of the present invention, the self-gating activation function ennSwish in step S3 dynamically adjusts the feature response, specifically:

[0035]

[0036] Among them, x is the input value, and the sigmoid function is used to map the input to between 0 and 1.

[0037] As another improvement of the present invention, in the region proposal network of step S4, feature extraction is performed on the input feature map through the shared convolutional layer of the RPN, and a series of candidate regions (Anchors) are generated using a sliding window. The features within each sliding window are mapped into a feature vector of a fixed dimension. Then, these feature vectors pass through two parallel fully connected branches to perform classification and bounding box regression operations on the generated feature vectors respectively. The classification branch is used to determine whether each Anchor contains an object (foreground or background), and the bounding box regression branch is used to perform precise position adjustment on the Anchor containing the object. Subsequently, all candidate regions are sorted according to the classification scores, and redundant candidate boxes with high overlap are removed through non-maximum suppression (NMS). Finally, a series of final high-quality target region candidate boxes are selected, and these candidate boxes are used as region proposals for potential targets.

[0038] As a further improvement of the present invention, the improved RoI detector in step S5 includes a classical RoIHead and three cascaded RiRoI Heads. First, the candidate boxes enter the classical RoI detector for coarse-grained screening to quickly complete preliminary object detection. Then, the screened candidate boxes are passed to three sequentially connected RiRoI Heads. The RiRoI Head introduces the RiRoI Align operation that can achieve both spatial alignment and orientation alignment, thereby ensuring the extraction of rotation-invariant features.

[0039] As a further improvement of the present invention, in order to further reduce the positioning error caused by rotation angle and scale changes, the KFIoU loss function is adopted in the bounding box regression process, replacing the traditional loss function, as shown in the following formula:

[0040] L KFIoU = Focal(KL(P pred , P gt )) = (1 - KL(P pred , P gt )) γ ·log(KL(P pred , P gt ))

[0041] Wherein, P pred and P gt respectively represent the probability distributions of the predicted bounding box and the ground truth bounding box. The KL divergence is used to accurately measure the difference between the two, and more precisely capture the overlapping degree of the rotated bounding boxes. Focal represents an improved cross-entropy loss function, and γ is a modulation factor used to control the weights of easy and hard samples. By reducing the weights of easy-to-classify samples and increasing the attention to hard samples, the problem of data imbalance is alleviated.

[0042] During the whole process, the detection results of each stage are gradually transmitted through a progressive optimization process of classification and bounding box regression, and finally a detection image labeled with high-precision target category and position information is output.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] (1) The improved ReDet framework combining multi-scale fusion and cascaded detection proposed by the present invention effectively improves the detection accuracy and robustness of the model for small rotated targets in complex backgrounds by optimizing feature extraction, feature fusion, and feature alignment. Compared with the traditional ReDet framework, the present invention not only ensures a high detection accuracy, but also effectively balances the model complexity and inference speed, and is particularly suitable for being deployed on resource-constrained UAV platforms, providing an efficient and reliable solution for intelligent power inspection.

[0045] (2) The present invention designs an improved rotation-equivariant recursive feature fusion pyramid - ReSiRFP, which significantly enhances the multi-scale and multi-modal feature modeling ability of the model for insulators and their defects in UAV aerial images, thereby improving the detection performance and effectively solving the problem of insufficient multi-scale feature fusion.

[0046] (3) The present invention designs a cascaded rotation-invariant RoI detector - CRiRoI Head, which significantly improves the feature expression ability of the detection model and significantly reduces the missed detection and false detection rates through a four-stage cascaded feature alignment mechanism, providing more reliable detection ability for insulator inspection in complex environments. Description of the Drawings

[0047] Figure 1 It is a schematic diagram of the step flow of the method of the present invention;

[0048] Figure 2 It is a schematic diagram of the structure of the improved feature pyramid module (ReSiRFP) in the method of the present invention;

[0049] Figure 3 It is a schematic diagram of the structure of the rotation-equivariant atrous spatial pyramid pooling module (ReSiASPP) in the method of the present invention;

[0050] Figure 4 It is a schematic diagram of the structure of the improved RoI detector (CRiRoI Head) in the method of the present invention;

[0051] Figure 5 It is a visual comparison diagram of the detection effects of several different methods in the test example of the present invention; among them,

[0052] Figure (a) shows the visualization diagram of the detection result of the R 3 det model;

[0053] Figure (b) shows the visualization diagram of the detection result of the Oriented RCNN model;

[0054] Figure (c) shows the visualization diagram of the detection result of the RoI Transformer model;

[0055] Figure (d) shows the visualization diagram of the detection result of the method of the present invention. Detailed Embodiments

[0056] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0057] Embodiment 1

[0058] An aerial insulator defect detection method based on multi-scale fusion and cascaded detection, as Figure 1 shown, the specific steps are as follows:

[0059] Step S1: Data preprocessing.

[0060] Image enhancement is performed on the publicly available datasets of drone aerial photography insulators (CPLID, InsPLAD) to expand the samples. Subsequently, annotation and data partitioning are carried out, and the improved ReDet model is loaded to complete the initialization.

[0061] The publicly available datasets of drone aerial photography insulators (CPLID, InsPLAD) are obtained from public data sources, and a variety of data augmentation techniques are used to expand the datasets, including specific operations such as random rotation, cropping, scaling to 145, and brightness adjustment, etc., to enhance data diversity and the generalization ability of the model. Then, the insulators and their defects in the datasets are accurately annotated. Subsequently, the annotated data is partitioned into a training set and a test set in a ratio of 9:1, which are used for model training and performance evaluation respectively.

[0062] Load the pre-trained ReDet model as the basic network and improve it by replacing the original feature pyramid module with the ReSiRFP module and replacing the RoI detector with the CRiRoI Head. The improved ReDet model includes at least a feature extraction module, an improved feature pyramid module (ReSiRFP),

[0063] a region proposal network, and an improved RoI detector (CRiRoI Head), and complete the initialization settings of the model,

[0064] providing a basic framework for subsequent training and detection.

[0065] Step S2: Feature extraction.

[0066] The input image is input into the backbone network 155 (ReResNet) of the improved model after the preprocessing operation in Step S1. The backbone network extracts the high-level features of the image and generates a feature map. ReResNet performs preliminary convolutional feature extraction on the input image, extracts features at a fixed number of rotation angles through rotation-equivariant convolution, and combines the weight sharing mechanism to ensure the consistency of features under rotation changes. Then, the features are further enhanced and fused in the residual block to generate a feature map with rotation equivariance.

[0067] First, the image is subjected to feature extraction through the rotation-equivariant convolutional layer. This process is based on the convolutional operation of the rotation-equivariant network and adopts a weight sharing mechanism in N directions, that is, each convolutional kernel is in C NShared parameters are used among all rotation transformations on the group to generate a feature map with rotation equivariance characteristics. Subsequently, the feature map undergoes a rotation-equivariant pooling operation (Pooling) to downsample in spatial resolution, ensuring the consistency of features in different directions. Then, the feature map is subjected to rotation-equivariant normalization, enabling balanced distribution of features in different directions. Finally, the rotation-equivariant feature map generated by ReResNet has a size of (K, N, H, W), where K is the number of channels, N is the number of direction channels, each direction channel corresponds to a rotation angle of the group C N and H and W are the spatial dimensions of the feature map.

[0068] Step S3: Obtaining multi-scale feature maps.

[0069] The feature map obtained in step S2 is fed into a rotation-equivariant recursive feature pyramid (ReSiRFP) to generate a feature map containing multi-scale information.

[0070] The feature map of the input image first enters an improved feature pyramid module (ReSiRFP), as specifically Figure 2 shown. Through bottom-up and top-down feature transfer and lateral connections, a preliminary multi-scale fused feature map is generated. Subsequently, the feature map enters a feedback loop, which introduces the ReSiASPP module, as Figure 3 shown. In this module, the input feature map first extracts local detail features through a common convolutional layer, and then passes through three dilated convolutional branches. The receptive field of the convolutional kernel in each branch is determined by its dilation rate: 1) Dilation rate of 6: used to capture local detail information of the input feature; 2) Dilation rate of 12: expand the receptive field to capture medium-scale context information; 3) Dilation rate of 18: further expand the receptive field for extracting global semantic information of large objects. At the same time, the feature map also passes through a global convolutional branch for global semantic information of large objects. In each branch, a rotation-equivariant gated self-controlled activation function ennSwish is used, and its core calculation formula is:

[0071]

[0072] where x is the input value, and the sigmoid function is used to map the input to between 0 and 1.

[0073] Subsequently, the feature maps generated by the three dilated convolution branches and the global convolution branch are concatenated in the channel dimension to integrate local details, context information of different receptive fields, and global semantic information, providing richer multi-scale features for recursive optimization. Then, these feature maps are fed back into the backbone network to re-extract new features. The multi-scale feature maps generated after all the above operations then enter the Fusion module in the output stage to complete dynamic feature fusion. Specifically, the feature maps are first processed by equivariant convolution (ennConv) to compress the number of feature channels to 8 equivariant channels, ensuring that the feature maps can maintain equivariance for 8 different rotation angles. Finally, the feature maps are fused by the Fusion module, and through global feature aggregation operations, combined with the Sigmoid activation function, a pixel-wise adaptive weight map W is generated to dynamically adjust the current-step feature map F curr and the previous iteration feature map F prev 's contribution ratio, and the calculation formula is:

[0074]

[0075] where F rot-equiv is the equivariant feature map, ennConv is the equivariant convolution operation, F multi-scale is the multi-scale feature map obtained through the Feature Pyramid Network, W is the weight map generated by F rot-equiv , σ is the sigmoid function, ennAdaptiveAvgPool is the equivariant average pooling operation, F fused is the output feature map after weighted fusion, F curr is the current-step feature, and F prev is the previous-step feature, and ⊙ represents element-wise multiplication.

[0076] Step S4: Region Proposal Network processing.

[0077] The feature maps obtained in step S3 are fed into the Region Proposal Network (RPN). The RPN slides on the feature maps to generate a series of candidate regions (Anchors), and then potential target region candidate boxes are screened out through classification and bounding box regression.

[0078] First, the RPN receives the feature map output from the ReResNet. Assume the dimension of the feature map is (W, H, C), where W and H are the width and height of the feature map, and C is the number of channels. Then, a set of reference boxes (Anchors) are generated at each pixel point of the feature map. Each pixel point is the center of an Anchor, and k = m·n Anchors are generated according to different scales and aspect ratios, where m represents the number of scales and n represents the number of aspect ratios. The coordinates of each Anchor are determined by the center point and the width and height. After generating the Anchors, the feature map extracts intermediate features through a shared convolutional layer and feeds them into two sub-networks: a classification sub-network and a bounding box regression sub-network. The classification sub-network is used to determine whether each Anchor contains an object; the bounding box regression sub-network is used to regress and correct the coordinates of the Anchors. Subsequently, the obtained Anchors are screened layer by layer: First, remove the Anchors that exceed the image boundary or have inappropriate sizes; then, classify the foreground and background according to the IoU (Intersection-over-Union) value between the Anchors and the ground truth boxes. Those with an IoU greater than 0.7 are marked as foreground, those less than 0.3 are marked as background, and the rest are ignored; finally, the remaining candidate regions are sorted according to the classification scores, and the candidate boxes with excessive overlap are removed using non-maximum suppression (NMS), and only the boxes with an IoU less than the threshold (such as 0.7) are retained. Finally, the RPN outputs a set of high-quality candidate regions, which include the coordinates of the candidate boxes and the corresponding classification scores for further processing by subsequent modules.

[0079] Step S5: Processing by the improved RoI detector.

[0080] The candidate boxes obtained in step S4 are fed into the improved RoI detector (CRiRoI Head) for spatial and orientation feature alignment, and the target category is predicted through the classification branch, and the position of the candidate boxes is adjusted through the bounding box regression branch, and finally a detection image labeled with multi-scale target categories and positions is generated.

[0081] To refine the candidate regions generated by the RPN and generate a detection image containing multi-scale target category and position information, such as Figure 4As shown, the candidate boxes are fed into the CRiRoI Head module and finely screened through cascaded detection to improve the detection performance. Specifically, first, the candidate boxes enter the classical RoI detector. First, the candidate boxes are spatially aligned through the RoI Align module, and the feature map is mapped into a region feature map of a fixed size using bilinear interpolation. Subsequently, these aligned feature maps undergo a feature expansion operation and are fed into the fully connected layer to complete the preliminary object classification and bounding box regression, generating the detection results of the first stage. Then, the detection results of the first stage are passed to three sequentially connected RiRoI Heads for refinement stage by stage. In each RiRoI Head, first, the RiRoI Align module is used to align the features of the rotated object, including two parts: spatial alignment and orientation alignment. The spatial alignment part extracts a region feature map of a fixed size from the rotation-equivariant feature map; the orientation alignment part calculates the orientation index according to the rotation angle, performs a cyclic shift on the orientation channel, and interpolates the angles not in the discrete rotation group to complete the smooth alignment of the orientation dimension. The aligned feature maps undergo a feature expansion operation and are then fed into the fully connected layer for classification and bounding box regression optimization.

[0082] In the four cascaded RoI detectors, the IoU threshold is increased stage by stage, and the positive and negative samples are reclassified and the bounding box positions are optimized. The detection results of each stage are used as the input of the next stage and are passed through the progressive optimization process of classification and boundary regression step by step to gradually improve the detection accuracy, and finally, a detection map labeled with high-precision object category and position information is output.

[0083] During the bounding box regression process, in order to better handle the positioning errors caused by rotated boxes and scale changes, an improved KFIoU (Kullback-Leibler Focal Intersection over Union) loss function is used to replace the traditional loss function. The basic form of KFIoU is:

[0084] L KFIoU = Focal(KL(P pred ,P gt ))=(1 - KL(P pred ,P gt )) γ ·log(KL(P pred ,P gt ))

[0085] Among them, P pred and P gtThey respectively represent the probability distributions of the predicted bounding boxes and the ground truth bounding boxes. The KL divergence is used to accurately measure the difference between the two, and more precisely capture the overlapping degree of the rotated bounding boxes. Focal represents an improved cross-entropy loss function, where γ is a modulation factor used to control the weights of easy and hard samples. By reducing the weights of easy-to-classify samples and increasing the attention to hard samples, the problem of data imbalance is alleviated.

[0086] The design of the entire cascaded detector further enhances the expression ability of small object features and the separation ability in complex environments, realizes the step-by-step refined screening of candidate regions, and finally generates detection results that accurately label the target category and location.

[0087] Test case

[0088] The experimental environment of this test case is as follows: GPU NVIDIA GeForce RTX 3080T, CPU Intel Core i7-11700KF@3.60GHz, Pytorch1.10, Ubuntu18.04, CUDA11.1.

[0089] To construct a high-quality dataset of transmission line insulator images, this test case combines the Chinese Power Line Insulator Dataset (CPLID) with the InsPLAD dataset collected by the Federal University of Pernambuco in Brazil, making full use of the rich insulator image resources of both. These image resources cover various types of insulators, including glass insulators, composite insulators, and their defect samples, such as defective glass insulators and defective composite insulators. On this basis, the original data is comprehensively processed through various data augmentation strategies, specifically including operations such as cropping, rotation, scaling, brightness adjustment, contrast adjustment, chromaticity change, and saturation change. Finally, a diverse self-built dataset containing 4116 images is generated. To ensure the accuracy and standardization of dataset annotation, the present invention uses the CVAT software to accurately annotate all images in the form of rotated bounding boxes (Oriented BoundingBox, OBB), marks the insulators and their defect regions in the form of rotated bounding boxes, and generates an annotation file containing detailed category labels and anchor box coordinate information. The number of category labels for each category is shown in Table 1.

[0090] Table 1 Statistical distribution of targets in the self-built dataset

[0091]

[0092] To verify the effectiveness of the method of the present invention in the task of UAV insulator defect detection, several commonly used rotated object detection methods are selected for comparative experiments with the method of the present invention, including the one-stage detection algorithm R 3Det, two-stage detection algorithms such as Oriented R-CNN, RoI Transformer, and the classic ReDet.

[0093] Figure 5 Visualized the detection effects of five algorithms on the self-built dataset. Among them, Figure (a) shows the visualization diagram of the detection results of the R 3 det model, Figure (b) shows the visualization diagram of the detection results of the Oriented RCNN model, Figure (c) shows the visualization diagram of the detection results of the RoI Transformer model, and Figure (d) shows the visualization diagram of the detection results of the method of the present invention. It can be seen from Figure (a) that in the multi-scale small target (polymer_insulator) localization task, there are different degrees of deviations in other algorithms, and for the insulator targets in the vertical form at a long distance, the first three algorithms cannot detect them, and the classic ReDet has a low localization accuracy. In contrast, the method of the present invention can accurately identify and precisely locate such targets. Figure (b) shows the detection performance under complex conditions such as high exposure and large-angle rotation. The method of the present invention is still superior to the other four comparison methods in terms of recognition and localization accuracy, and further improves the detection effect while maintaining high accuracy; Figure (c) shows the significant advantage of the method of the present invention in the case of target occlusion and complex background. Especially when detecting severely occluded defective targets (glass defect), other algorithms are easily interfered by the occluder, resulting in a large difference in the detection accuracy of occluded and unoccluded defective targets, while the method of this paper can maintain a high detection accuracy in both cases; The scenario shown in Figure (d) contains multiple complex features such as multiple targets (glass_defect, defective_glass_insulator, ploymer_insulator), small targets, large-angle rotation, strong occlusion, and poor lighting. In this scenario, the detection effect of the method of the present invention is still comprehensively superior to other algorithms.

[0094] Table 2 shows the comprehensive performance of five algorithms on the self-built dataset, including the average detection accuracy of normal insulators, defective insulators, and defects (denoted as AP1, AP2, and AP3 respectively), the overall average detection accuracy (mAP), the number of parameters of the algorithm model (Params), and the number of frames of images processed per second by the algorithm (FPS).

[0095] Table 2 Comparison of experimental results of different algorithms

[0096]

[0097] As can be seen from the above table, the method proposed in the present invention performs best among the five algorithms, with an overall mAP reaching 94.02%. Compared with the classical ReDet before improvement, the overall mAP of the method of the present invention has increased by 3.36%. Among them, the detection accuracy of defective targets has been improved most significantly, reaching 2.27%; the detection accuracies of normal insulators and defective insulators have increased by 1.72% and 1.98% respectively. At the same time, the number of model parameters and the calculation speed have not changed much compared with classical ReDet, fully demonstrating the advantage that the method of the present invention significantly improves the detection accuracy while maintaining light weight.

[0098] Therefore, the present invention introduces an improved feature pyramid module (ReSiRFP), effectively enhancing the rotational equivariance of the model and the multi-scale feature integration ability; at the same time, an improved RoI detector (CRiRoI head) is constructed by combining a four-stage cascade strategy and a new directional loss function KFIoU, further improving the model's ability to express target features. Generally speaking, this method significantly reduces the missed detection and false detection rates in complex backgrounds, and significantly improves the detection accuracy and robustness of small rotating targets in complex backgrounds.

[0099] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.

Claims

1. Aerial photography insulator defect detection method based on multi-scale fusion and cascade detection, characterized by: The steps include: S1. Data preprocessing: preprocess the data in the drone aerial photography insulator dataset, and annotate and divide the preprocessed dataset before entering the improved ReDet model; the improved ReDet model at least includes a feature extraction module, an improved feature pyramid module, a region proposal network and an improved RoI detector; S2, feature extraction: the data preprocessed in step S1 is fed into the improved ReDet model, feature extraction is performed in the feature extraction module, and a feature map with rotational equivariance is generated; S3, multi-scale feature map acquisition: input the rotation equivariant feature map generated in step S2 into an improved feature pyramid module to obtain a multi-scale feature map; the improved feature pyramid module at least includes a rotation equivariant recursive feature pyramid network, a rotation equivariant void space pyramid pooling module and a Fusion module; S4, region proposal network processing: sending the multi-scale feature map obtained in step S3 to the region proposal network for processing, wherein the region proposal network at least includes a shared convolutional layer, a classification branch and a boundary regression branch, wherein the classification branch and the boundary regression branch are two parallel fully connected branches; A sliding window mechanism is used to slide on the feature map to generate a series of candidate regions, and potential target region candidate frames are screened out through classification and boundary regression operations; S5, improved RoI detector processing: input the target area candidate box obtained in step S4 into the improved RoI detector, align the spatial and directional features, and then predict the target category through the classification branch and adjust the candidate box position through the boundary regression branch to generate a detection image containing the target category and position information.

2. The method for detecting insulator defects by aerial photography based on multi-scale fusion and cascade detection according to claim 1, characterized in that: The data preprocessing in step S1 at least includes random rotation, cropping, scaling and brightness adjustment. After the data set is labeled, the data set is divided into a training set and a test set in a ratio of 9:

1.

3. The method for detecting insulator defects by aerial photography based on multi-scale fusion and cascade detection according to claim 1, characterized in that: In the feature extraction of step S2, the feature extraction module performs preliminary feature extraction on the input data through rotational equivariant convolution, and generates a feature map with rotational equivariance through rotational equivariant pooling and rotational equivariant normalization under a multi-directional weight sharing mechanism; the size of the generated rotational equivariant feature map is (K, N, H, W), where K is the number of channels, N is the number of directional channels, and H and W are the spatial sizes of the feature map.

4. The method for detecting insulator defects by aerial photography based on multi-scale fusion and cascade detection according to claim 1, characterized in that: The step S3 of acquiring the multi-scale feature map specifically includes the following steps: S31: The input feature map enters the rotation equivariant recursive feature pyramid network, and generates a preliminary multi-scale feature map through bottom-up and top-down feature transfer; S32: Preliminary multi-scale feature map. The rotation-equivariant atrous spatial pyramid pooling module in the feedback link of the improved feature pyramid module is used to extract rotation-equivariant multi-scale features through atrous convolutions with different atrous rates, and the feature response is dynamically adjusted through the self-gated activation function ennSwish. S33: After the multi-scale feature map is once again transmitted from bottom to top and from top to bottom, it passes through the Fusion module in the output link of the improved feature pyramid module, dynamically adjusts the contribution of each scale feature through an adaptive weight mechanism, and outputs the final multi-scale feature map; the use of the adaptive weight mechanism to dynamically adjust the contribution of each scale feature is specifically as follows: Among them, F rot-equiv is the rotation equivariant feature map, ennConv is the rotation equivariant convolution operation, and F multi-scale is the multi-scale feature map obtained through the feature pyramid network, and W is the rot-equiv The generated weight map, σ is the sigmoid function, ennAdaptiveAvgPool is the rotation equivariant average pooling operation, F fused is the feature map output after weighted fusion, F curr is the current step feature and F prev is the feature of the previous step, and ⊙ represents the element-wise dot product.

5. The method for detecting insulator defects by aerial photography based on multi-scale fusion and cascade detection according to claim 4, characterized in that: The self-gating activation function ennSwish in step S3 dynamically adjusts the characteristic response, specifically: Among them, x is the input value, and the sigmoid function is used to map the input to between 0 and 1.

6. The method for detecting insulator defects by aerial photography based on multi-scale fusion and cascade detection according to claim 1, characterized in that: In the region proposal network of step S4, a shared convolutional layer extracts features from the input feature map, and a series of candidate regions are generated using a sliding window. The features in each sliding window are mapped into a feature vector of a fixed dimension, and classification and boundary regression operations are performed; wherein the classification branch is used to determine whether each candidate region contains a target, and the boundary regression branch is used to accurately adjust the position of the candidate region containing the target; all candidate regions are sorted according to the classification scores, and redundant candidate frames with high overlap are removed by non-maximum suppression, and the final target region candidate frame is screened out as the region frame of the potential target.

7. The method for detecting insulator defects by aerial photography based on multi-scale fusion and cascade detection according to claim 1, characterized in that: The improved RoI detector in step S5 includes a cascade of a classic RoIHead and three RiRoI Heads. The candidate box enters the classic RoIHead for coarse-grained screening, and the screened candidate box is passed to three sequentially connected RiRoI Heads for spatial alignment and directional alignment RiRoI Align operations; wherein the classification and boundary regression results are transmitted in a step-by-step optimization manner.

8. The method for detecting insulator defects by aerial photography based on multi-scale fusion and cascade detection according to claim 7, characterized in that: In the boundary regression of step S5, the KFIoU loss function is used, specifically: L KFIoU =Focal(KL(P pred ,P gt ))=(1-KL(P pred ,P gt )) γ ·log(KL(P pred ,P gt )) Among them, Focal is an improved cross entropy loss function, KL divergence is used to accurately measure the difference between the two, P pred and P gt They represent the probability distribution of the predicted box and the true box respectively, and γ is the adjustment factor.

Citation Information

Cited By

  • Low-proportion infrared small target detection method and device

    CN121767790A