Insulator Defect Detection Method and System Improved Based on RT-DETR
By improving the insulator defect detection method of RT-DETR, using StarNet star network, partial self-attention mechanism and improved convolution module, combined with Inner-IoU and CIoU loss functions, the problems of backbone network redundancy and computing complexity in the prior art are solved, and more efficient and accurate insulator defect detection are achieved.
Patent Information
- Application Number
- CN202411874576.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-19
AI Technical Summary
In the existing insulator defect detection technology, the backbone network has redundancy in feature extraction, ADD operations increase computational complexity, and attention mechanism increases network complexity, resulting in low detection efficiency and poor generalization capabilities.
The insulator defect detection method of RT-DETR is improved, the StarNet star network is used as the backbone network, the P5 detection head is replaced as the P2 detection head, the star convolution module is improved, the multi-head attention mechanism is used to replace the partial self-attention mechanism (PSA), the RepC3 convolution module is improved as the RepMob convolution module, and the Inner-IoU and CIoU loss functions are combined.
Through the improved network structure and modules, redundancy in feature extraction is reduced, computational complexity is reduced, detection efficiency and accuracy is improved, and the generalization ability of the model is enhanced.
Smart Images

Figure CN119323572B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of insulator defect detection, and particularly to an insulator defect detection method and system improved based on RT-DETR. Background Art
[0002] Insulators are an important part of transmission lines and are prone to damage when exposed to outdoor conditions for a long time. Effectively detecting insulator defects has important practical significance. In traditional methods, insulator defect detection relies on manual inspection. However, with the emergence of deep learning in this field, object detection algorithms have gradually replaced traditional manual inspection. Compared with the costly and time-consuming manual inspection, object detection algorithms can efficiently and safely complete the insulator defect detection task.
[0003] Currently, object detection algorithms are generally divided into single-stage detection algorithms and two-stage detection algorithms. Although the two-stage detection algorithms have a higher accuracy ceiling, they are accompanied by higher and more complex computing costs, and the cost of non-maximum suppression in post-processing is too high, making it impossible to achieve end-to-end object detection, which is not conducive to real-time object detection and not suitable for detecting insulator defects; while the single-stage detection algorithms require a process, have lower computing costs, and the detection process is relatively simple. The YOLO series is a very commonly used single-stage detection algorithm, but it still relies on non-maximum suppression to obtain a unique prediction box during the inference stage.
[0004] In the invention with the patent application number CN202410495935.5, a lightweight insulator defect detection method based on improved YOLOv8 is proposed, and the training is mainly carried out according to the following steps: (1) Collect the insulator fault dataset for power transmission line inspection, and preprocess and augment the collected insulator images; (2) Obtain the improved YOLOv8 network model and train the network model using the processed dataset; (3) Input the insulator data to be detected into the optimal detection model, and output the insulator defect detection results, that is, the fault category of the insulator and the location where the defect appears. However, this solution has the following defects: ① In the improvement of the backbone network in this solution, a large number of residual networks are used for feature fusion, which will make the network more complex. In the block of FasterNet used, the ADD operation is used for feature fusion. This fusion method will increase the redundancy of features, and the ADD operation will increase the computational complexity and result in low performance. ② The EMA attention mechanism and the SPPF module used in the backbone network of this solution not only have similar functions, causing redundancy and repeated extraction of multi-scale information, but also increase the complexity of the backbone network, which is not conducive to the subsequent input of the transformer module of RT-DETR. ③ In the GELAN module of the efficient layer aggregation network used in this solution, due to aggregating the information of many branches and then performing overall gradient update, the amount of information calculated is extremely large and the computational complexity is high. ④ In the inference stage of this solution, many redundant prediction boxes will be generated, relying on NMS to process the backend data, which is not conducive to the end-to-end detection task, will make the inference efficiency low, and is not conducive to the real-time detection of insulators.
[0005] In the invention with the patent application number CN202410326526.2, a traffic sign detection method based on improved RT-DETR is proposed, and the training is mainly carried out according to the following steps: 1) Obtain public traffic sign images and divide the target detection data sets required for training and testing; 2) Design a traffic sign detection model of improved RT-DETR, and propose the MCCA (Multi channelcoordinated attention) attention mechanism; To improve the generation ability of offset and mask parameters, the MCCA attention module is spliced in the DCNv2 deformable convolution to generate the DCNv2att module; Two BasicBlock blocks in the detection network Backbone are replaced with DCNv2att blocks; 3) Evaluate the training results of the traffic sign detection model based on improved RT-DETR. However, this solution has the following defects: ① In this method, too many residual connection branches are used in the latter part of the backbone network, which will greatly increase the complexity of the network, and the efficiency of splicing features using the ADD operation is low, and this traditional convolutional network has limitations in the high-dimensional non-linear transformation of feature representation. ② If the data set deviation is large, the parameters learned by the deformable convolution will be more targeted at a certain data set, and the detection effect will deteriorate and the generalization ability will be low when switching data sets. ③ In the MCCA attention mechanism module of this method, aggregating features along the two spatial directions of x and y is redundant. In the basic module of RT-DETR, the convolutional network, AIFI, and CCFM can all aggregate spatial information well. Adding this attention mechanism will make the network structure redundant, and due to the addition of two branches, the calculation of gradient information will be more complex.
[0006] In the invention with the patent application number CN202410005858.0, a method and system for detecting infrared small and weak aircraft by RT-DETR are proposed, and the training is mainly carried out according to the following steps: 1) Select an infrared small and weak aircraft data set and divide it into a training set, a test set and a validation set; 2) Based on the RT-DETR network, change the Backbone backbone network to the VanillaNet minimalist neural network; 3) Replace the RepC3 structure in the neck of the RT-DETR network with the DWRC3 structure; 4) Introduce three partial convolution layers at the end of the neck of the RT-DETR network as the 28th, 29th, and 30th layers, and use them as the outputs of the 12th, 15th, and 18th layers of DWRC3 respectively to complete the construction of the initial detection model; 5) Use the training set and the validation set to train the constructed initial detection model to obtain the trained detection model; 6) Use the trained detection model to detect the target to be detected. However, this solution has the following defects: ① The VanillaNet used in this method is relatively weak in extracting some useful regions due to its relatively low non-linearity, and has weak ability to fuse multiple types of information due to extremely few network branches. ② In the improved DWRC3 of this method, the SR part is replaced, and ordinary convolution layers are used, resulting in low feature extraction efficiency. In addition, a large number of residual structures are added, and the ADD operation is used to fuse features, which will cause a lot of redundancy, computational complexity and parameter quantity in the network; ③ Using the traditional P3, P4, P5 detection heads without optimizing for small components is not conducive to defect detection of small targets. ④ The loss function used does not consider the direction of mismatch between the required ground truth boxes and the predicted boxes, resulting in a slow convergence speed and low efficiency, because the predicted boxes may shift during training and ultimately lead to a worse model effect.
[0007] The disadvantages of the above prior art can be summarized as follows: For the backbone network among them, the network structure will cause redundancy of feature map information when extracting features of images, and widely using the ADD operation to fuse feature maps will lead to low fusion efficiency and greatly increase the network width, making the network complex and not concise enough. The traditional convolutional network has limitations in the high-dimensional non-linear transformation of feature representation; the improved attention mechanism has many additional branches, which is very complex when calculating gradient information and will reduce the inference speed; the RepC3 module of the original network has no additional branches, which will consume a certain amount of computational resources compared with using branches, and the convolutional module is not efficient enough. Using this kind of convolution for inference is not conducive to improving the accuracy and cannot replace the convolutional network with branches. Moreover, after improving this module, the excessive additional branches have caused complex gradient information update and network complexity. In addition, the bounding box loss function among them cannot self-adjust according to different detectors and detection tasks, does not have strong generalization, and ignores the influence of the matching angle between the ground truth boxes and the predicted boxes. Summary of the Invention
[0008] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide an insulator defect detection method and system improved based on RT-DETR.
[0009] To achieve the above purpose, the present invention provides the following technical solutions: An insulator defect detection method improved based on RT-DETR, the detection method includes the following steps:
[0010] Collect the publicly available insulator defect dataset and obtain images by intercepting frames from videos as the dataset, and split the dataset into a training set and a validation set.
[0011] Improve based on the object detection model RT-DETR, change the backbone network to the StarNet star network for feature extraction, further improve the StarNet star network, remove the P5 detection head, add a P2 detection head, and improve the star convolution module in the StarNet star network.
[0012] Change the multi-head attention mechanism in RT-DETR to the partial self-attention mechanism PSA in YOLOv10, and further improve the partial self-attention mechanism PSA to obtain the S-PSA attention mechanism module.
[0013] Improve the RepC3 convolution module in RT-DETR to the RepMob convolution module.
[0014] Combine the loss function with Inner-IoU and CIoU to obtain the Inner-CIoU loss function for loss calculation.
[0015] Input the training set into the improved RT-DETR for training, and verify it on the validation set to obtain the best weight on the validation set.
[0016] Load the obtained best weight into the improved RT-DETR for insulator defect detection.
[0017] In a preferred embodiment, collecting the publicly available insulator defect dataset and obtaining images by intercepting frames from videos as the dataset includes the following steps:
[0018] First, collect the publicly available insulator defect dataset, then collect the drone inspection video data provided by the power grid, obtain the foreign object dataset by intercepting frames from the video data, and use the LabelImg tool to label the foreign objects in the images, indicating the category and the true bounding box.
[0019] In a preferred embodiment, improving the star convolution module in the StarNet star network includes the following steps:
[0020] Input feature information is split into two channels through a split layer. One channel is upsampled through point convolution, and the other channel goes through depth convolution and is then split in half and passed through fully connected layers respectively. Half of them use the ReLU6 activation function to make the channels non-linear, and the other half remains unchanged. Then, the features of the two are fused through a star operation, and the features of different channels are multiplied pairwise. After that, depth convolution is used to integrate and strengthen the fused features. Then, the channel is concatenated with the channel of the point convolution in terms of channel dimension, and then a channel shuffle operation is performed to enhance the communication between channels. Finally, it is integrated through a normal convolution layer.
[0021] In a preferred embodiment, the processing flow of the S-PSA attention mechanism module is as follows: First, the input information is subjected to feature extraction through depthwise separable convolution. Subsequently, the obtained feature information passes through the ReLU6 activation function to introduce non-linearity into the features. Then, the feature information is split through a split layer into two parts, 1 / 4 and 3 / 4. The 1 / 4 feature information remains unchanged. 1 / 3 of the 3 / 4 feature information remains unchanged, and 2 / 3 of the 3 / 4 feature information undergoes a multi-head self-attention mechanism operation. Then, it is fused with 1 / 3 of the 3 / 4 feature information that remains unchanged through a star operation. Then, the feature information is split in half. Half of it goes through point convolution to increase the dimension and fuse partial channel information, and then dilated convolution is used to further obtain global information. The result after convolution is combined with the other half of the feature information that remains unchanged through a star operation to infiltrate the learned global information into the channels. Then, it is concatenated with the original 1 / 4 feature information, and then channel shuffle is performed to enable information exchange between the learned global information and the original 1 / 4 feature information. Finally, it is integrated through normal convolution.
[0022] In a preferred embodiment, the RepC3 convolution module in RT-DETR is improved to a RepMob convolution module, including the following steps:
[0023] During the training phase, the input information is divided into three parts and passed through three branches. The first branch does not process the information. The second branch uses a 3×3 depth convolution to extract features. The third branch uses a 1×1 depth convolution to extract features. Then, the feature information of the three branches is fused through a star operation. Next, the SE channel attention mechanism is used to enhance the channel information and distinguish the importance of the channel information. Subsequently, the channel information is split into two parts. One part undergoes point convolution to change the dimension information, followed by dimensionality reduction, then passes through the ReLU6 activation function, and then undergoes point convolution to adjust the channel dimension and increase the dimension. Finally, a star operation is performed with the other part of the channel information that is not processed to obtain the final output. During the inference phase, the initial three branches are merged into a 3×3 depth convolution. Then, the SE channel attention mechanism module calculates the weights for the channels to obtain important and unimportant channel information. The channel information is also split into two parts. One part is not processed, and the other part is dimensionally increased through point convolution for the channel information. After passing through the ReLU6 activation function, a star operation is performed with the unprocessed part of the channel information to obtain the output.
[0024] In a preferred embodiment, the loss function is combined with Inner-IoU and CIoU to obtain the Inner-CIoU loss function for loss calculation, including the following steps:
[0025] The calculation method of Inner-IoU is as follows:
[0026] , ,
[0027] , ,
[0028] ,
[0029] , ,
[0030] where is the left boundary of the ground truth box, is the right boundary of the ground truth box, is the bottom boundary of the ground truth box, is the top boundary of the ground truth box, is the left boundary of the predicted box, is the right boundary of the predicted box, is the bottom boundary of the predicted box, is the top boundary of the predicted box, is the abscissa of the center point of the ground truth box, is the ordinate of the center point of the ground truth box, is the abscissa of the center point of the predicted box, is the vertical coordinate of the center point of the prediction box, is the width of the ground truth box, is the height of the ground truth box, w is the width of the prediction box, h is the height of the prediction box, ratio is the scale factor, inter is the intersection of the auxiliary box of the ground truth box and the auxiliary box of the prediction box, union is the union of the auxiliary box of the ground truth box and the auxiliary box of the prediction box, and finally obtain is the IoU of the auxiliary box, and then combine it with the loss of CIoU, , , , , ,
[0031] where b is the center coordinates of the prediction box, is the center point coordinates of the ground truth box, is to calculate the Euclidean distance, is the weight function, v is the parameter for measuring the similarity of the aspect ratio, c is the diagonal distance of the minimum bounding rectangle, is the ratio of the intersection area of the ground truth box and the prediction box to the union area, is the new IoU obtained by considering the complete intersection between the target boxes and introducing a correction factor, is the loss of CIoU, is the loss of Inner-IoU combined with CIoU.
[0032] The insulator defect detection system improved based on RT-DETR includes a data preprocessing module, an improved backbone network module, an improved attention mechanism module, an improved RepC3 convolution module, an improved loss function module, a model training module, and an insulator defect detection module;
[0033] Data preprocessing module: Collect the publicly available insulator defect dataset and obtain images by intercepting frames from videos as the dataset, and split the dataset into a training set and a validation set;
[0034] Improved backbone network module: Based on the object detection model RT-DETR, improve it by changing the backbone network to the StarNet star network for feature extraction, further improve the StarNet star network, remove the P5 detection head, add a P2 detection head, and improve the star convolution module in the StarNet star network;
[0035] Improved attention mechanism module: Change the multi-head attention mechanism in RT-DETR to the partial self-attention mechanism PSA in YOLOv10, and further improve the partial self-attention mechanism PSA to obtain the S-PSA attention mechanism module;
[0036] Improve the RepC3 Convolution Module: Improve the RepC3 convolution module in RT-DETR to the RepMob convolution module;
[0037] Improve the Loss Function Module: Combine the loss function with Inner-IoU and CIoU to obtain the Inner-CIoU loss function for loss calculation;
[0038] Model Training Module: Input the training set into the improved RT-DETR for training and verify it on the validation set to obtain the best weights on the validation set;
[0039] Insulator Defect Detection Module: Load the best weights into the improved RT-DETR for insulator defect detection.
[0040] In the above technical solution, the technical effects and advantages provided by the present invention are as follows:
[0041] 1. The present invention uses the StarNet star network as a new backbone network for improvement. Through star operations, while maintaining a relatively low computational complexity, it realizes the mapping of high-dimensional feature spaces, achieves richer feature representations through star operations, and the present invention further improves the StarNet star network by adding a P2 detection head to make small object detection more accurate and improving the star convolution module to make the feature extraction part more efficient;
[0042] 2. The present invention improves the multi-head attention mechanism module by performing partial self-attention operations and then merging them, without performing attention operations on all information, which will not cause waste of computational resources of the attention mechanism and can also make good use of the attention mechanism to pay more attention to defects; the present invention also improves the RepC3 convolution module by using the self-developed RepMob convolution module, adding a certain number of branches without increasing excessive computational amounts, and also improving the convolution module by using star operations to improve the information fusion part, making the top-down and bottom-up feature fusions more efficient, improving the utilization rate of feature information. The present invention also combines Inner-IoU and CIoU, uses auxiliary bounding boxes to improve the calculation of the bounding box loss function, and considering the center point distance and diagonal distance of the bounding box, uses a combined function of the two for bounding box loss calculation to improve the model convergence speed and training accuracy. Description of the Drawings
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0044] Figure 1 is the method flowchart of the present invention;
[0045] Figure 2 is the schematic diagram of the improved StarNet star network;
[0046] Figure 3 is the schematic diagram of the improved star convolution module;
[0047] Figure 4 is the schematic diagram of the S-PSA attention mechanism module;
[0048] Figure 5 is the schematic diagram of the training stage of the RepMob convolution module;
[0049] Figure 6 is the schematic diagram of the inference stage of the RepMob convolution module. Specific embodiments
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0051] Embodiment 1: Please refer to Figure 1 As shown, the insulator defect detection method improved based on RT-DETR in this embodiment includes the following steps:
[0052] (1) Collect the publicly available insulator defect dataset and obtain images by intercepting frames from videos as the dataset;
[0053] First, collect the publicly available insulator defect dataset, and then collect the drone inspection video data provided by the power grid. For the video data, obtain the foreign object dataset by intercepting frames to obtain images, and use the LabelImg tool to label the foreign objects in the images, indicating the categories and the real boxes. Split the dataset into a training set and a validation set (8:2).
[0054] (2) Improve based on the object detection model RT-DETR. Change the backbone network to the StarNet star network for feature extraction. Further improve the StarNet star network by removing the P5 detection head and adding the P2 detection head to improve the small object detection effect, and improve the star convolution module in the StarNet star network to make the features obtained by the network more abundant and efficient;
[0055] The improved StarNet star network is as follows Figure 2 As shown, the image is input into the StarNet star network. The convolutional downsampling module is used to first obtain the features with 2x downsampling, and then continue with the convolutional downsampling operation to obtain the features with 4x downsampling. Then, the features are fused and enriched through the star convolutional module, and then the image features with 8x and 16x downsampling are continuously obtained. The features with 32x downsampling are removed because the insulator itself is relatively small, and the 32x downsampling detection head of P5 is generally used to detect large targets. In this embodiment, the P5 detection head is discarded, and the 4x downsampling detection head of P2 is added, which is more conducive to the detection of small targets such as insulators.
[0056] The improved star convolutional module is as follows Figure 3 As shown, first, the feature information is input. The channel is split through the split layer into two channels. One channel uses point convolution for convolutional upsampling to improve cross-channel information fusion. The other channel undergoes depth convolution. Since depth convolution performs convolution on each channel separately and uses different convolutional kernels, the extracted features can also have diversity, strengthening the feature representation ability of the network. After depth convolution, it is divided into two halves and passed through the fully connected layer respectively. One half uses the ReLU6 activation function to make the channel non-linear, and the other half is not processed. Then, through the star operation, the features of the two are fused, and the features of different channels are multiplied pairwise, realizing efficient feature representation. Without excessive additional computational overhead, it can perform calculations in the low-dimensional space and can implicitly consider extremely high-dimensional features, providing efficient calculations and better feature representation. Then, through depth convolution, the fused features are integrated and strengthened. Then, this channel is concatenated with the channel of point convolution in the channel dimension. Then, because depth convolution lacks information exchange between its own channels and also lacks communication with the channels of point convolution, a channel shuffle operation is added to strengthen the communication between channels. Finally, it is integrated through the ordinary convolutional layer. The improved star convolutional module has simple calculations and is more efficient for feature extraction.
[0057] The present invention uses the StarNet star network to improve the backbone network of the original model and further improves the StarNet star network, making the feature extraction part of the model more efficient and with simple operations.
[0058] (3) Change the multi-head attention mechanism in RT-DETR to the partial self-attention mechanism PSA in YOLOv10, and further improve the partial self-attention mechanism PSA to obtain the S-PSA attention mechanism module. Use the partial self-attention mechanism to perform self-attention operations on part of the information, and convolve the other information and fuse it with the information after self-attention operations, which can also achieve the effect of self-attention and can also reduce resource waste;
[0059] As shown in Figure 4 , the processing flow of the S-PSA attention mechanism module is as follows: First, the input information is subjected to depthwise separable convolution for feature extraction, which can generate diverse feature information and amplify and extract important features. The purpose is to better associate favorable information globally when performing the multi-head attention mechanism calculation subsequently. Then, the obtained feature information passes through the ReLU6 activation function, and this computationally simple non-linear activation function introduces non-linearity into the features, thereby introducing it into the model to make the model non-linear. Then, the feature information is split by the split layer into two parts, 1 / 4 and 3 / 4. The 1 / 4 feature information remains unchanged, 1 / 3 of the 3 / 4 feature information remains unchanged, and 2 / 3 of the 3 / 4 feature information undergoes the multi-head self-attention mechanism operation. In this way, half of the original information undergoes the multi-head self-attention operation to capture global feature information. Then, it is fused with 1 / 3 of the 3 / 4 feature information that remains unchanged using the star operation. Using the star operation instead of the ADD operation for information fusion can achieve high computational efficiency while obtaining richer and more expressive feature representations. Then, the feature information is split into two halves. One half passes through point convolution to increase the dimension and fuse partial channel information, and then through dilated convolution to further obtain more global information. To prevent the grid effect caused by dilated convolution, the result after its convolution is combined with the other half of the feature information that remains unchanged using the star operation, infiltrating the learned global information into the channels and improving the information loss of dilated convolution. Then, it is concatenated with the original 1 / 4 feature information, and then channel shuffle is performed to enable information exchange between the learned global information and the original 1 / 4 feature information, that is, both the original information is retained and better global information is learned. Finally, it is integrated to the appropriate number of channels through ordinary convolution and output to the next layer to continue processing the information.
[0060] The present invention uses a partial self-attention mechanism module to improve the multi-head self-attention mechanism of the original method, enabling the network to capture global features without overly increasing the computational complexity brought by the multi-head self-attention mechanism, making the acquisition of global information more efficient.
[0061] (4) Improve the RepC3 convolution module in RT-DETR to a RepMob convolution module;
[0062] This embodiment refers to the reparameterization structure of RepVIT and improves the RepC3 convolution module to a RepMob convolution module. During the training phase, as shown in Figure 5As shown in the figure, first, the input information is divided into three parts and passed through three branches. The first branch does not process the information, aiming to retain the original information. The second branch uses a 3×3 depth convolution to extract features, and the third branch uses a 1×1 depth convolution to extract features. Image information with different receptive field sizes is obtained using convolution kernels of different sizes. Then, the feature information of the three branches is fused through a star operation. Then, the SE channel attention mechanism is used to strengthen the channel information, distinguish the importance of the channel information, and enable the model to focus on more important channel information. Subsequently, the channel information is split into two parts. One part changes the dimension information through point convolution, performs dimensionality reduction, then passes through the ReLU6 activation function, and then adjusts the channel dimension through point convolution to increase the dimension, enabling the learning of more channel dimension information. Finally, a star operation is performed on it and the other part of the channel information that is not processed to obtain the final output. And during the inference stage after reparameterization, as Figure 6 shown, the initial three branches are merged into a 3×3 depth convolution, and then the SE channel attention mechanism module calculates the weights for the channels to obtain important and unimportant channel information. The channel information is also split into two parts. One part is not processed, and the other part increases the dimension of the channel information through point convolution, passes through the ReLU6 activation function, and then performs a star operation with the unprocessed part of the channel information to obtain the output.
[0063] The present invention improves the RepC3 convolution module. During the training stage, depth convolution and the SE channel attention mechanism are used to continuously amplify the channel features and focus on more important channel information, enabling the network to learn more distinguishable features during training. During inference, the number of branches is reduced, and the convolution branches are merged into a single depth convolution with high computational efficiency. Moreover, the SE channel attention mechanism and the star operation are used to greatly improve the retention of features during the training stage, enabling a simple inference network structure to achieve the effect during training.
[0064] (5) Combine the loss function with Inner-IoU and CIoU to obtain the Inner-CIoU loss function for loss calculation;
[0065] Inner-IoU is a method for calculating the loss by using an auxiliary bounding box. It controls the size of the auxiliary bounding box through a scale factor ratio to calculate the loss and accelerate convergence. The calculation method of Inner-IoU is as follows:
[0066] , ,
[0067] , ,
[0068] ,
[0069] , ,
[0070] where is the left boundary of the ground truth box, is the right boundary of the ground truth box, is the bottom boundary of the ground truth box, is the top boundary of the ground truth box, is the left boundary of the predicted box, is the right boundary of the predicted box, is the bottom boundary of the predicted box, is the top boundary of the predicted box, is the abscissa of the center point of the ground truth box, is the ordinate of the center point of the ground truth box, is the abscissa of the center point of the predicted box, is the ordinate of the center point of the predicted box, is the width of the ground truth box, is the height of the ground truth box, w is the width of the predicted box, h is the height of the predicted box, ratio is the scale factor, inter is the intersection of the auxiliary box of the ground truth box and the auxiliary box of the predicted box, union is the union of the auxiliary box of the ground truth box and the auxiliary box of the predicted box, and finally we get is the IoU of the auxiliary box, and then it is combined with the loss of CIoU, , , , , ,
[0071] where b is the center coordinates of the predicted box, is the center point coordinates of the ground truth box, is to calculate the Euclidean distance, is the weight function, v is the parameter for measuring the similarity of the aspect ratio, c is the diagonal distance of the minimum bounding rectangle, is the ratio of the intersection area to the union area of the ground truth box and the predicted box (Intersection over Union), is the new IoU obtained by considering the complete intersection between the target boxes and introducing a correction factor, is the loss of CIoU, is the loss of Inner-IoU combined with CIoU.
[0072] The present invention uses Inner-IoU combined with CIoU to calculate the loss, accelerates the loss convergence by using the auxiliary bounding box, and at the same time takes into account the center point distance and diagonal distance of the bounding box, making the loss calculation more reasonable and effective, and accelerating the convergence speed of the model.
[0073] (6) Input the training set into the improved RT-DETR for training, and validate it on the validation set to obtain the best weights on the validation set;
[0074] (7) Load the obtained best weights into the improved RT-DETR for insulator defect detection.
[0075] Example 2: The insulator defect detection system improved based on RT-DETR described in this example includes a data preprocessing module, an improved backbone network module, an improved attention mechanism module, an improved RepC3 convolution module, an improved loss function module, a model training module, and an insulator defect detection module;
[0076] Data preprocessing module: Collect the publicly available insulator defect dataset and obtain images by intercepting frames from videos as the dataset, and split the dataset into a training set and a validation set;
[0077] Improved backbone network module: Improve based on the object detection model RT-DETR, change the backbone network to the StarNet star network for feature extraction, further improve the StarNet star network, remove the P5 detection head, add a P2 detection head, and improve the star convolution module in the StarNet star network;
[0078] Improved attention mechanism module: Change the multi-head attention mechanism in RT-DETR to the partial self-attention mechanism PSA in YOLOv10, and further improve the partial self-attention mechanism PSA to obtain the S-PSA attention mechanism module;
[0079] Improved RepC3 convolution module: Improve the RepC3 convolution module in RT-DETR to the RepMob convolution module;
[0080] Improved loss function module: Combine the loss function with Inner-IoU and CIoU to obtain the Inner-CIoU loss function for loss calculation;
[0081] Model training module: Input the training set into the improved RT-DETR for training, and validate it on the validation set to obtain the best weights on the validation set;
[0082] Insulator defect detection module: Load the obtained best weights into the improved RT-DETR for insulator defect detection.
[0083] In the description of this specification, the descriptions referring to the terms "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in a suitable manner in any one or more embodiments or examples.
[0084] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to only the specific embodiments. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. The improved insulator defect detection method based on RT-DETR is characterized by: The detection method comprises the following steps: Collect publicly available insulator defect datasets and obtain images by capturing video frames as datasets, and split the datasets into training sets and validation sets; Based on the target detection model RT-DETR, the backbone network is changed to StarNet network for feature extraction, and the StarNet network is further improved by removing the P5 detection head and adding the P2 detection head, and improving the convolution module in the StarNet network; The multi-head attention mechanism in RT-DETR is changed to the partial self-attention mechanism PSA in YOLOv10, and the partial self-attention mechanism PSA is further improved to obtain the S-PSA attention mechanism module; Improve the RepC3 convolutional module in RT-DETR to the RepMob convolutional module; The loss function is combined with Inner-IoU and CIoU to obtain the Inner-CIoU loss function for loss calculation; The training set is input into the improved RT-DETR for training, and then verified in the validation set to obtain the best weight in the validation set; The best weights are loaded into the improved RT-DETR for insulator defect detection; Improving the star convolution module in the StarNet network includes the following steps: Input feature information, split the channel into two through the split layer, one channel uses point convolution for dimension increase, and the other channel undergoes deep convolution, and then is divided into two halves and passes through the fully connected layer respectively. One half uses the ReLU6 activation function to make the channel nonlinear, and the other half is not processed. Then the features of the two are fused through star operation, the features of different channels are multiplied two by two, and the fused features are integrated and strengthened through deep convolution. Then the channel is concat-spliced with the channel of point convolution in the channel dimension, and then the channel shuffle operation is performed to strengthen the communication between channels, and finally integrated through the ordinary convolution layer.
2. The improved insulator defect detection method based on RT-DETR according to claim 1 is characterized in that: Collecting a public insulator defect dataset and obtaining images by capturing video frames as a dataset includes the following steps: First, we collect the public insulator defect dataset, and then collect the drone inspection video data provided by the power grid. We use the method of capturing frames from the video data to obtain the foreign object dataset. We use the LabelImg tool to annotate the foreign objects in the image, indicating the category and the true frame.
3. The improved insulator defect detection method based on RT-DETR according to claim 2 is characterized in that: The processing flow of the S-PSA attention mechanism module is as follows: first, the input information is subjected to deep separable convolution for feature extraction, and then the obtained feature information is subjected to the ReLU6 activation function to introduce nonlinearity into the feature, and then the feature information is split through the split layer into two parts, 1 / 4 and 3 / 4, 1 / 4 feature information is not operated, 1 / 3 of the 3 / 4 feature information is not operated, and 2 / 3 of the 3 / 4 feature information is operated by the multi-head self-attention mechanism, and then it is fused with the 1 / 3 of the 3 / 4 feature information that is not operated by star operation, and then the feature information is split into two halves, one half is increased in dimension by point convolution, and part of the channel information is fused, and then the global information is further obtained by the void convolution, and the convolution result is combined with the other half of the feature information that is not operated by star operation, the learned global information is infiltrated into the channel, and then concat spliced with the original 1 / 4 feature information, and then the learned global information is exchanged with the original 1 / 4 feature information through channel shuffling, and finally integrated through ordinary convolution.
4. The improved insulator defect detection method based on RT-DETR according to claim 3 is characterized in that: The RepC3 convolutional module in RT-DETR is improved to the RepMob convolutional module, which includes the following steps: In the training phase, the input information is divided into three parts and then passed through three branches. The first branch does not process the information, the second branch uses 3×3 deep convolution to extract features, and the third branch uses 1×1 deep convolution to extract features. Then, the feature information of the three branches is fused through star operation, and then the SE channel attention mechanism is used to strengthen the channel information and distinguish the importance of the channel information. The channel information is then split into two parts. One part changes the dimensional information through point convolution for dimensionality reduction, and then passes through the ReLU6 activation function, and then adjusts the channel dimension through point convolution to increase the dimension. Finally, it is star-operated with the other part of the channel information that is not processed to obtain the final output; in the inference phase, the initial three branches are merged into a 3×3 deep convolution, and then the channel weights are calculated through the SE channel attention mechanism module to obtain important channels and unimportant channel information. The channel information is also split into two parts, one part is not processed, and the other part is channel information increased in dimension through point convolution, and then after the ReLU6 activation function, a star operation is performed with the unprocessed part of the channel information to obtain the output.
5. The improved insulator defect detection method based on RT-DETR according to claim 4 is characterized in that: The loss function is combined with Inner-IoU and CIoU to obtain the Inner-CIoU loss function for loss calculation, including the following steps: Inner-IoU is calculated as follows: , , , , , , , in is the left boundary of the real box, is the right boundary of the real box, is the bottom boundary of the real box, is the top boundary of the real box, is the left boundary of the prediction box, is the right boundary of the prediction box, is the bottom boundary of the prediction box, is the top boundary of the prediction box, is the horizontal coordinate of the center point of the real frame, is the ordinate of the center point of the real frame, is the horizontal coordinate of the center point of the prediction box, is the ordinate of the center point of the prediction box, is the width of the real frame, is the height of the real box, w is the width of the predicted box, h is the height of the predicted box, ratio is the scale factor, inter is the intersection of the auxiliary box of the real box and the auxiliary box of the predicted box, and union is the union of the auxiliary box of the real box and the auxiliary box of the predicted box. Finally, we get is the IoU of the auxiliary box, which is then combined with the loss of CIoU. , , , , , Where b is the center coordinate of the prediction box, is the coordinate of the center point of the real frame, To calculate the Euclidean distance, is the weight function, v is a parameter to measure the similarity of aspect ratios, c is the diagonal distance of the minimum enclosing rectangle, It is the ratio of the intersection area of the real box and the predicted box to the combined area. In order to consider the complete intersection between target boxes and introduce the new IoU after the correction factor, is the CIoU loss, It is the loss of Inner-IoU combined with CIoU.
6. An improved insulator defect detection system based on RT-DETR, used to implement the detection method according to any one of claims 1 to 5, characterized in that: Including data preprocessing module, improved backbone network module, improved attention mechanism module, improved RepC3 convolution module, improved loss function module, model training module, insulator defect detection module; Data preprocessing module: collects publicly available insulator defect datasets and obtains images by capturing video frames as datasets, and splits the datasets into training sets and validation sets; Improved backbone network module: Based on the target detection model RT-DETR, the backbone network is changed to StarNet network for feature extraction, and the StarNet network is further improved by removing the P5 detection head and adding the P2 detection head, and improving the convolution module in the StarNet network; Improved attention mechanism module: The multi-head attention mechanism in RT-DETR is changed to the partial self-attention mechanism PSA in YOLOv10, and the partial self-attention mechanism PSA is further improved to obtain the S-PSA attention mechanism module; Improved RepC3 convolution module: The RepC3 convolution module in RT-DETR is improved to RepMob convolution module; Improved loss function module: The loss function is combined with Inner-IoU and CIoU to obtain the Inner-CIoU loss function for loss calculation; Model training module: input the training set into the improved RT-DETR for training, and verify it in the validation set to obtain the best weight in the validation set; Insulator defect detection module: The optimal weight is loaded into the improved RT-DETR for insulator defect detection.
Citation Information
Patent Citations
RT-DETR infrared weak aircraft detection method and system
CN117974972A
Traffic sign detection method based on improved RT-DETR
CN117975418A
Lightweight insulator defect detection method based on improved YOLOv8
CN118314436A
Lightweight improvement-based L-YOLOv8s insulator defect detection method
CN119006427A
Lightweight fire dynamic monitoring method and system
CN119007110A