A method for identifying typical defects in transmission lines based on binocular images
By using a binocular feature fusion module, feature screening module and improved FCOS target detection head in the transmission line defect detection, the misidentification and misidentification problems caused by image interference in the prior art are solved, and higher detection accuracy and small target positioning capabilities are achieved.
Patent Information
- Application Number
- CN202211556482.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-06
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-12-06
AI Technical Summary
The existing transmission line defect detection methods are susceptible to interference from irrelevant information in the image, resulting in misidentification and misidentification.
The transmission line defect detection method based on binocular image is adopted. The binocular feature fusion module makes full use of binocular image information, and the feature filtering module filters irrelevant information, and improves the FCOS target detection head to improve the positioning ability of small targets.
It improves the accuracy of defect detection in transmission lines, reduces interference from complex backgrounds on defect identification, and enhances the positioning ability of small targets such as broken strands and foreign objects.
Smart Images

Figure CN116309270B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and relates to a method for identifying typical defects of transmission lines based on binocular images. Background Art
[0002] Due to being exposed to the outside for a long time, transmission lines sometimes have defects such as broken strands and foreign object suspension. Identifying these defects in the drone inspection images is a key task for subsequent defect handling by staff to eliminate potential safety hazards.
[0003] The resolution of transmission line images is usually high, but the sizes of transmission line defects such as broken strands and foreign objects are small, especially the broken strand defects, which are difficult to detect even by the human eye. It is difficult to guarantee the accuracy of traditional methods based on digital image processing. Some researchers have applied object detection methods based on deep learning to the task of identifying transmission line defects, but these methods also have some deficiencies. They are easily interfered by irrelevant information in the images, and it is difficult to accurately identify and locate small-scale transmission line defects, resulting in false recognition and missed recognition. Summary of the Invention
[0004] The present invention provides a method for identifying typical defects of transmission lines based on binocular images to solve the problem that existing transmission line defect detection methods are easily interfered by irrelevant information in the images, resulting in more false recognition and missed recognition.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A method for detecting transmission line defects based on binocular images specifically includes the following steps:
[0007] Step 1: Obtain a binocular transmission line image dataset, perform distortion correction and epipolar correction on the images in the dataset to obtain the corrected left and right images;
[0008] Step 2: Construct a network model for detecting transmission line defects based on binocular images. The network model includes two identical backbone networks, four binocular feature fusion modules with different scales, a feature screening module, and an improved FCOS object detection head, where:
[0009] The two backbone networks are respectively used to extract multi-scale features of the left and right images. Specifically, the left and right images are respectively input into two parallel backbone networks to obtain four left feature maps and right feature maps Their length and width sizes are (1 / 4, 1 / 8, 1 / 16, 1 / 32) of the original image size;
[0010] The binocular feature fusion module is used to input the four pairs of left and right feature maps obtained by the two backbone networks into the binocular feature fusion modules of corresponding scales respectively for fusion, so as to obtain binocular fusion feature maps at four different resolutions.
[0011] The feature screening module is used to filter the binocular fusion feature maps, and then construct a binocular fusion feature pyramid with four levels through a top-down path.
[0012] The improved FCOS object detection head is used to make predictions on the feature maps of four scales in the binocular fusion feature pyramid respectively, and calculate the final detection results of transmission line defects.
[0013] The improved FCOS object detection head has three outputs, namely classification prediction, bounding box prediction and IoU score prediction. The classification prediction branch outputs the class confidence scores, that is, the probabilities of each sample point belonging to various classes. The bounding box prediction branch is used to output the bounding box, that is, the distances of the sample point from the four sides of the predicted bounding box of the target object. The IoU prediction branch is used to output the IoU prediction scores corresponding to the bounding box prediction.
[0014] Step 3: Train the transmission line defect detection model based on binocular images on the binocular transmission line defect dataset.
[0015] Step 4: Input the binocular transmission line image to be recognized into the trained network model for defect detection. When making predictions, multiply the class confidence score of each sample point output by the network by the corresponding IoU prediction score to obtain a new bounding box detection confidence score. Among them, the class confidence score is the prediction result obtained by the classification prediction branch, the IoU prediction score is the result obtained by the IoU score prediction branch, and the bounding box detection confidence score is the result obtained by supervising the bounding box output by the bounding box prediction branch using the IoU Loss. The IoU Loss is a loss function that directly optimizes the bounding box measurement index.
[0016] Then use the new bounding box detection confidence score to perform non-maximum suppression operation, and finally obtain the final defect detection result.
[0017] Furthermore, the binocular feature fusion module in Step 2 is used to implement the following process:
[0018] First, input the left feature map and the right feature map of the same scale into the binocular similarity calculation module to calculate the binocular similarity matrix to indirectly represent the depth information of the binocular images, and use a 3×3 convolutional layer to filter and integrate the binocular similarity matrix to obtain an implicit depth feature map; then use 1×1 convolutions to adjust the number of channels for the current left and right feature maps respectively; then concatenate the feature maps of the left and right images after channel adjustment and the implicit depth feature map along the channel direction, and input the concatenated feature map into the ASPP module, and finally obtain the binocular fusion feature map.
[0019] Further, the binocular similarity calculation module is used to implement the following steps:
[0020] A. First, keep the left feature map unchanged, move the right feature map d units to the right, and then perform a dot product operation on the corresponding channel vectors at the overlapping positions of the left and right feature maps to obtain a two-dimensional similarity map at an offset of d, whose size is w×h, representing the similarity between each pixel point on the left feature map and the pixel point with an offset of d on the right feature map;
[0021] B. Perform the same operation as in step A on the left and right feature maps at different offsets, and then stack the two-dimensional similarity maps calculated at these different offsets together along the channel direction to obtain a three-dimensional binocular similarity matrix M of size w×h×d max where d max represents the maximum offset.
[0022] Further, the calculation of the binocular similarity matrix is shown in the following formula:
[0023]
[0024] where M represents the binocular similarity matrix, d represents the offset, c represents the number of channels of the left feature map or the right feature map, F l represents the left feature map, F r represents the right feature map, and x and y respectively represent the ordinate and abscissa on the feature map.
[0025] Further, the ASPP module implements the following functions: perform convolution operations with a 3×3 convolutional kernel by three dilation convolutional layers with different dilation rates, where the dilation rates are 1, 2, and 4 respectively, and then concatenate them, and pass the concatenated feature map through a 1×1 convolutional layer.
[0026] Further, in step 2, the feature screening module is used to implement the following process:
[0027] The binocular fusion feature map is input into the channel attention module to calculate the attention weight vector, which is then multiplied by the feature map to obtain the feature map after weight adjustment. Then, through skip connection, the feature map after weight adjustment is added to the input feature map to output the binocular fusion feature pyramid of different scales after feature screening based on residual attention.
[0028] Further, the channel attention module is used to implement the following functions:
[0029] First, through two-dimensional adaptive average pooling and max pooling operations, the spatial information of the input feature map is aggregated to obtain two vectors with the same length as the number of channels of the original feature map. Then, these two vectors are non-linearly transformed and added, and then interval mapping is performed through the sigmoid function to finally obtain the channel attention weight vector.
[0030] Further, in step 3, the loss function used for training is as follows:
[0031] L total = L cls + L loc + L IoU
[0032] Among them, L loc is the calibration box regression loss, L cls is the class confidence loss, and L IoU is the IoU score prediction loss.
[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0034] (1) The present invention designs a binocular feature fusion module, which makes more full use of the information of binocular images and improves the accuracy of transmission line defect detection.
[0035] (2) The present invention filters out irrelevant information in the feature map through the feature screening module, making the attention of the network more concentrated on which the target object is, and reducing the interference of complex backgrounds on defect recognition.
[0036] (3) The present invention improves the FCOS detection head, modifies the centerness prediction branch to the IoU prediction branch, and improves its positioning ability for small targets such as broken strands and foreign objects. Brief Description of the Drawings
[0037] Figure 1 is the structural diagram of the transmission line defect detection network based on binocular images;
[0038] Figure 2 is the structural diagram of the binocular fusion module;
[0039] Figure 3Schematic diagram of the calculation process of the binocular similarity matrix;
[0040] Figure 4 It is the ASPP module structure diagram;
[0041] Figure 5 It is the feature screening module structure diagram;
[0042] Figure 6 It is the schematic diagram of the channel attention mechanism;
[0043] Figure 7 It is the structure diagram of the improved FCOS object detection head;
[0044] Figure 8 They are the foreign object image detection result (a) and the broken strand image detection result (b) images. Specific implementation manner
[0045] The present invention will be described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0046] The transmission line defect detection method based on binocular images given by the present invention includes the following steps:
[0047] Step 1: Obtain a binocular transmission line image dataset, perform distortion correction and epipolar correction on the pictures in the dataset to obtain the corrected left and right images.
[0048] In this embodiment, the binocular transmission line image dataset contains 1,781 pairs of transmission line binocular pictures. Some pictures in the dataset were taken at the power supply bureau, and the other part was obtained by building a simulated transmission line scene and using a drone equipped with a ZED binocular camera for shooting. All binocular images in the dataset were subjected to distortion correction and epipolar correction using the remap function of OpenCV. The LabelImg tool was used to perform object detection annotation on the dataset. The number of pictures of different types in the dataset is shown in Table 1. The two types of pictures containing broken strands and foreign objects account for the majority, and there are also a small number of normal line pictures and a small number of transmission line pictures containing both broken strand and foreign object defects.
[0049] Table 1 Number of pictures of different categories
[0050]
[0051] Step 2: Construct a transmission line defect detection network model based on binocular images. The network model includes two identical backbone networks, four binocular feature fusion modules with different scales, a feature screening module, and an improved FCOS object detection head, where:
[0052] Two backbone networks are respectively used to extract multi-scale features of the left and right images. Specifically, the left and right images are respectively input into two parallel backbone networks to obtain four left feature maps at different scales and right feature maps Their length and width dimensions are (1 / 4, 1 / 8, 1 / 16, 1 / 32) of the original image size;
[0053] The binocular feature fusion module is used to input the four pairs of left and right feature maps obtained by the two backbone networks into the binocular feature fusion module at the corresponding scales for fusion, and obtain binocular fusion feature maps at four different resolutions
[0054] The feature screening module is used to filter the binocular fusion feature maps, and then construct a binocular fusion feature pyramid with four levels through a top-down path
[0055] The improved FCOS object detection head is used to make predictions on the feature maps at four scales in the binocular fusion feature pyramid respectively, and calculate the final transmission line defect detection result.
[0056] Among them, as Figure 2 shown, the detailed calculation process of the binocular feature fusion module in step 2 is as follows:
[0057] First, the left and right feature maps of the same scale are input into the Binocular Similarity Module (BSM) to mine the correlation between the pixels on the same epipolar line of the left and right feature maps, calculate the binocular similarity matrix to indirectly represent the depth information of the binocular images, and use a 3×3 convolutional layer to filter and integrate the binocular similarity matrix to obtain an implicit depth feature map. Then, the current left and right feature maps are respectively adjusted in the number of channels using 1×1 convolutions to reduce the subsequent computation amount of the binocular fusion module. Then, the feature maps of the left and right images after the channel number adjustment and the implicit depth feature map are concatenated along the channel direction, and the concatenated feature map is input into the Atrous Spatial Pyramid Pooling (ASPP) module, and finally the binocular fusion feature map is obtained.
[0058] The calculation process of the described binocular similarity calculation module BSM is as Figure 3As shown, where the light gray box represents the left feature map and the dark gray box represents the right feature map. The steps are as follows: A. First, keep the left feature map stationary, move the right feature map d units to the right, and then perform a dot product operation on the channel vectors at the corresponding positions in the overlapping part of the left and right feature maps to obtain a two-dimensional similarity map at an offset of d, with a size of w×h, representing the similarity between each pixel point on the left feature map and the pixel point with an offset of d on the right feature map. B. Perform the same operation on the left and right feature maps at different offsets as in step A, and then stack the two-dimensional similarity maps calculated at these different offsets along the channel direction to obtain a three-dimensional binocular similarity matrix M with a size of w×h×d max where d max represents the maximum offset.
[0059] The binocular similarity matrix described represents the similarity between the corresponding feature points of the left and right feature maps at different offsets. The larger its value, the smaller the difference between the two feature points. The calculation of the binocular similarity matrix is shown as follows:
[0060]
[0061] where M represents the binocular similarity matrix, d represents the offset, c represents the number of channels of the left or right feature map, F l represents the left feature map, F r represents the right feature map, and x and y represent the vertical and horizontal coordinates on the feature map respectively. The structure of the ASPP module is as Figure 4 shown. Convolution operations with a convolution kernel of 3×3 are performed by three dilated convolutional layers with different dilation rates, where the dilation rates are 1, 2, and 4 respectively, and then the three are concatenated. Finally, the concatenated feature map passes through a 1×1 convolutional layer.
[0062] The detailed calculation process of the feature screening module in step 2 is as follows: First, input the binocular fusion feature map calculated by the binocular feature fusion module into the channel attention module to calculate the attention weight vector, and multiply it with the feature map to obtain the feature map adjusted by the weight. Then, add the feature map adjusted by the weight to the input feature map through a skip connection to output four different-scale binocular fusion feature pyramids after feature screening based on residual attention. In this module, combining the attention mechanism with the residual connection can improve the adaptability of the feature screening module, and the skip connection can also improve the efficiency of gradient backpropagation during network training. The adjustment of the feature map weight by the feature screening module is shown as follows:
[0063] F′ = (1 + A(F)) × F
[0064] Among them, F′ represents the binocular fusion feature map after weight adjustment, F represents the input binocular fusion feature map, A represents the attention module, and A(F) represents the attention weight distribution tensor.
[0065] The structure of the channel attention module is as Figure 6 shown. First, through two-dimensional adaptive average pooling and max pooling operations, the spatial information of the input feature map is aggregated to obtain two vectors with the same length as the number of channels of the original feature map. Then, these two vectors are non-linearly transformed and added together, and then interval mapping is performed through the sigmoid function to finally obtain the channel attention weight vector. The calculation process of the channel attention mechanism is shown in the following formula:
[0066]
[0067] where σ represents the sigmoid function, MLP represents the fully connected layer, F represents the input feature map, W1 is the parameter vector of the first layer of the fully connected layer, and W2 is the parameter vector of the second layer of the fully connected layer. is the vector after the feature map passes through max pooling; is the vector after the feature map passes through average pooling.
[0068] The structure of the improved FCOS object detection head in step 2 is as Figure 7 shown. This detection head has three outputs, namely classification prediction, bounding box prediction, and IoU score prediction. The classification prediction branch outputs the probability that each sample point belongs to each category. The bounding box prediction branch outputs a four-dimensional vector for each sample point, indicating the distance of this sample point from the four sides of the predicted bounding box of the target object.
[0069] The improved FCOS object detection head modifies the centerness prediction branch in the original FCOS detection head into an IoU prediction branch. Although the centerness branch in the original FCOS detection head can suppress some low-quality prediction boxes that deviate from the center of the target object during the test phase, it is not friendly to the prediction of small targets. This centerness measurement method only considers the normalized distance between the sample point and the center point of the object. For small targets, as long as the distance between the sample point and the center point of the object is a few pixels larger, it may cause the value of centerness to decrease sharply. This makes the centerness prediction branch unstable during training and difficult to converge. In addition, the centerness branch may also lead to the missed detection of small target objects during the test phase because for some objects with very small sizes, it may cause the score predicted by the centerness branch to be very small, and thus the final prediction confidence also decreases accordingly. Using the IoU prediction branch instead of the centerness branch can alleviate these problems because centerness only considers the distance between the sample point and the center of the ground truth bounding box, while the IoU measurement method considers the overlap degree between the predicted bounding box and the ground truth bounding box, which is more conducive to the detection of small targets.
[0070] The described IoU prediction branch shares four convolutional layers with the bounding box regression branch, and finally uses a separate convolutional layer to predict the IoU score of each sample point in the feature map.
[0071] Step 3: Train the transmission line defect detection model based on binocular images on the binocular transmission line defect dataset.
[0072] During training, the loss function of the transmission line defect detection network based on binocular images consists of three parts, namely the bounding box regression loss L loc , the class confidence loss L cls and the IoU score prediction loss L IoU . The definition of the overall loss function is shown as follows:
[0073] L total = L cls + L loc + L IoU
[0074] Among them, the loss function L IoU of the IoU score prediction branch is shown as follows:
[0075]
[0076]
[0077] Where Pos is the set of positive sample points in the feature map, NPos is the number of positive sample points, N is the number of all sample points in the feature map, BCE is the binary cross-entropy loss function, and IoU i is the IoU score predicted by the IoU prediction branch for the i-th sample point in the feature map, is the IoU value between the predicted bounding box and the ground truth bounding box at the i-th sample point, represents the ground truth bounding box, and box i represents the predicted bounding box. intersection represents the area of the overlapping part of the two boxes, and union represents the area of the union of the two boxes.
[0078] The class confidence loss L cls is defined as follows:
[0079]
[0080] where p i is the probability value predicted by the algorithm that this sample point is a positive sample, and its range is between 0 and 1, represents the true class to which the i-th sample point belongs, and its value is 0 or 1. FL represents the focal loss function (Focal Loss).
[0081] For the bounding box prediction branch, IoU Loss is used for supervision. IoU Loss is a loss function that directly optimizes the bounding box measurement index. Compared with loss functions based on absolute distance such as L1 Loss and L2 Loss, it is more reasonable. It will not cause a large change in the order of magnitude of the loss function value due to the change in the scale of the target object, and is more suitable for the detection task of small targets such as transmission line defects. The definition of the bounding box regression loss is as follows:
[0082]
[0083] The batch size during training is 4, the learning rate is set to 0.0004, the optimizer is the adaptive optimizer Adam, and data augmentation is performed during training. Random cropping is synchronously performed on the same positions of the left and right images, and at least one defect will be retained on the cropped image when cropping the image containing defects.
[0084] Step 4: Input the binocular transmission line image to be identified into the trained network model for defect detection. When making predictions, multiply the category confidence score of each sample point output by the network by the corresponding IoU prediction score to obtain a new calibration box detection confidence score; the improved FCOS target detection head has three outputs, namely classification prediction, calibration box prediction and IoU score prediction; the classification prediction branch outputs the category confidence score, that is, the probability that each sample point belongs to each category, and the calibration box prediction branch is used to output the calibration box, that is, the distance between the sample point and the four sides of the predicted calibration box of the target object; the IoU score prediction branch is used to output the IoU prediction score corresponding to the calibration box prediction;
[0085] Then, the new calibration box detection confidence score is used to perform non-maximum suppression operations to filter out some prediction boxes that have large position deviations from the target object, and finally obtain the final defect detection result.
[0086] The defect detection results of some pictures are as follows Figure 8 It can be seen that the size of the transmission line defects is usually very small, especially the broken strand defects, which are difficult to distinguish even for human eyes, while the method proposed in the present invention successfully detects these defects.
[0087] The power transmission line defect detection method based on binocular images proposed in the present invention is compared with some commonly used target detection methods, and the experimental results are shown in Table 2. The compared methods include (1) Faster-RCNN: a two-stage target detection method. This method first uses a region generation network to obtain candidate regions where the target object may exist, and then inputs the feature map corresponding to the candidate region into a small classification subnetwork to obtain the final category and calibration frame coordinates. (2) Cascade R-CNN: a two-stage target detection method. This method adds multiple cascaded detectors with different thresholds on the basis of Faster-RCNN. (3) YOLOV4: the fourth version of the YOLO series target detection network. Its backbone network adopts the CSPDarkNet-53 network that takes both speed and accuracy into consideration, and uses the PANet module to improve the efficiency of information flow between feature maps. (4) RetinaNet: a single-stage anchor-based target detection method. This method alleviates the problem of category imbalance in target detection tasks by dynamically adjusting the weights of difficult and easy samples. (5) FCOS: a single-stage target detection method without anchor boxes.
[0088] Table 2 Experimental results of different methods
[0089]
[0090] As can be seen from the above experimental results, the transmission line defect detection method proposed by the present invention based on binocular feature fusion and improved FCOS detection head has achieved good results. The AP of broken strands reaches 87.78%, and the AP of foreign objects reaches 93.92%. The accuracy rate exceeds these typical object detection methods listed in the table. The best method other than the method in this paper is FCOS, and the mAP of the method in this paper is 4.14% higher than that of FCOS. In addition, in the experimental results of these methods, the AP of broken strands is lower than that of foreign objects, because the broken strand defects in the transmission line are finer, the features are less obvious, and it is more difficult to identify.
[0091] The present invention conducts experiments to compare the influence of different binocular fusion methods on the results, and compares three fusion methods respectively. Fusion method 1: left feature map + right feature map + depth feature fusion; Fusion method 2: left feature map + depth feature fusion; Fusion method 3: left feature map + right feature map fusion. The comparison experimental results are shown in Table 3.
[0092] Table 3 Comparison of three binocular fusion methods
[0093]
[0094] The experimental results show that among these three fusion methods, the best one is fusion method 1, that is, fusing both the left and right image features and the implicit depth features output by the BSM module. This method makes more full use of the information in the binocular images, and the recognition accuracy has been greatly improved compared with the case without binocular fusion.
[0095] The present invention conducts experiments to compare the effects of different attention mechanisms, and conducts comparative experiments on several different feature screening methods such as no feature screening, channel residual attention, spatial residual attention, and hybrid residual attention mechanisms. The experimental results are shown in Table 4:
[0096] Table 4 Comparison of different feature screening methods
[0097]
[0098] As can be seen from the table, the feature screening module based on channel residual attention has achieved the best results, because the channel attention mechanism can recalibrate the weights of each channel of the binocular fusion features, suppress the useless feature channels, and improve the response of the target object in the feature map. After adding the spatial residual attention module to the network, there is almost no improvement compared with the case of no feature screening, and even the AP of foreign objects has a slight decrease. It is speculated that because the spatial attention mechanism is more sensitive to the scene changes in the image, it is difficult to accurately place the attention weights on the spatial positions of the target objects.
[0099] To verify the effectiveness of the improved FCOS detection head, this paper conducted comparative experiments on the detection head before and after the improvement. The experimental results are shown in Table 5. It can be seen from the experimental results that the detection accuracy of the improved detection head for both types of defects, namely broken strands and foreign objects, has been improved. This shows that changing the centerness branch in FCOS to the IoU prediction branch has a better effect on small-sized objects such as transmission line defects.
[0100] Table 5 Experimental comparison results of improved FCOS
[0101]
Claims
1. A method for detecting transmission line defects based on binocular images, characterized in that Specifically, it includes the following steps: Step 1: Obtain a binocular transmission line image dataset, perform distortion correction and epipolar correction on the images in the dataset to obtain the corrected left and right images; Step 2: Construct a transmission line defect detection network model based on binocular images. The network model includes two identical backbone networks, four binocular feature fusion modules with different scales, a feature screening module, and an improved FCOS object detection head, where: The two backbone networks are respectively used to extract multi-scale features of the left and right images. Specifically, the left and right images are respectively input into two parallel backbone networks to obtain left feature maps at four different scales and right feature maps Their length and width dimensions are (1 / 4, 1 / 8, 1 / 16, 1 / 32) of the original image size; The binocular feature fusion module is used to input the four pairs of left and right feature maps obtained by the two backbone networks into the binocular feature fusion modules of corresponding scales for fusion respectively, so as to obtain binocular fusion feature maps at four different resolutions. The feature screening module is used to filter the binocular fusion feature map, and then construct a binocular fusion feature pyramid with four levels through a top-down path The improved FCOS object detection head is used to perform predictions on the feature maps of four scales in the binocular fusion feature pyramid respectively, and calculate the final transmission line defect detection result; the improved FCOS object detection head has a total of three outputs, namely classification prediction, bounding box prediction, and IoU score prediction; among them, the classification prediction branch outputs the class confidence score, that is, the probability that each sample point belongs to each category, the bounding box prediction branch is used to output the bounding box, that is, the distance from the sample point to the four sides of the predicted bounding box of the target object; the IoU score prediction branch is used to output the IoU prediction score corresponding to the bounding box prediction; Step 3: Train the transmission line defect detection model based on binocular images on the binocular transmission line defect dataset; Step 4: Input the binocular transmission line image to be recognized into the trained network model for defect detection. When making predictions, multiply the class confidence score of each sample point output by the network by the corresponding IoU prediction score to obtain a new bounding box detection confidence score; among them, the class confidence score is the prediction result obtained by the classification prediction branch, the IoU prediction score is the result obtained by the IoU score prediction branch, and the bounding box detection confidence score is the result obtained by supervising the bounding box output by the bounding box prediction branch using IoU Loss, where IoU Loss is a loss function that directly optimizes the bounding box measurement index; Then use the new bounding box detection confidence score to perform non-maximum suppression operation, and finally obtain the final defect detection result.
2. The method for detecting transmission line defects based on binocular images according to claim 1, wherein, The binocular feature fusion module in Step 2 is used to implement the following process: First, input the left and right feature maps of the same scale into the binocular similarity calculation module, calculate the binocular similarity matrix to indirectly represent the depth information of the binocular images, and use a 3×3 convolutional layer to filter and integrate the binocular similarity matrix to obtain an implicit depth feature map; then use 1×1 convolution to adjust the number of channels of the current left and right feature maps respectively; Then, concatenate the feature maps of the left and right images and the implicit depth feature map after adjusting the number of channels along the channel direction, and input the concatenated feature map into the ASPP module to finally obtain the binocular fusion feature map.
3. The method for detecting transmission line defects based on binocular images according to claim 2, wherein, The binocular similarity calculation module is used to implement the following steps: A. First, keep the left feature map unchanged, move the right feature map d units to the right, and then perform a dot product operation on the corresponding position channel vectors of the overlapping part of the left and right feature maps to obtain a two-dimensional similarity map at the offset of d, whose size is w×h, indicating the similarity between each pixel point on the left feature map and the pixel point with an offset of d on the right feature map; B. Perform the same operations on the left and right feature maps at different offsets as in step A, and then stack the two-dimensional similarity maps calculated at these different offsets together along the channel direction to obtain a three-dimensional binocular similarity matrix M of size w×h×d max , where d max represents the maximum offset.
4. The method for detecting transmission line defects based on binocular images according to claim 2, characterized in that The calculation of the binocular similarity matrix is shown as follows: Where M represents the binocular similarity matrix, d represents the offset, c represents the number of channels of the left feature map or the right feature map, F l represents the left feature map, F r represents the right feature map, and x and y represent the vertical and horizontal coordinates on the feature map respectively.
5. The method for detecting transmission line defects based on binocular images according to claim 2, wherein The ASPP module implements the following functions: perform convolution operations with a 3×3 convolution kernel by three dilated convolutional layers with different dilation rates, where the dilation rates are 1, 2, and 4 respectively, then splice the three of them, and pass the spliced feature map through a 1×1 convolutional layer.
6. The method for detecting transmission line defects based on binocular images according to claim 1, wherein In step 2, the feature screening module is used to implement the following process: Input the binocular fusion feature map into the channel attention module, calculate the attention weight vector, multiply it with the feature map to obtain the feature map adjusted by weights, then add the feature map adjusted by weights and the input feature map through skip connection, and output the binocular fusion feature pyramid of different scales after feature screening based on residual attention.
7. The method for detecting transmission line defects based on binocular images according to claim 6, characterized in that, The channel attention module is used to implement the following functions: First, through two-dimensional adaptive average pooling and max pooling operations, aggregate the spatial information of the input feature map to obtain two vectors with the same length as the number of channels of the original feature map, then perform non-linear transformation on these two vectors and add them, and then perform interval mapping through the sigmoid function to finally obtain the channel attention weight vector.
8. The method for detecting transmission line defects based on binocular images according to claim 1, wherein, In step 3, the loss function used for training is as follows: L total = L cls + L loc + L IoU Among them, L loc is the regression loss of the calibration box, L cls is the confidence loss of the category, and L IoU is the prediction loss of the IoU score.