Multi-object Detection Method, Device, Equipment and Medium for Infrared Images of Transmission Lines
The improved YOLOv5s algorithm with global context attention and context enhancement modules enhances detection precision and robustness for multiple targets in infrared images of power transmission lines, overcoming challenges posed by complex backgrounds and varying target sizes.
Patent Information
- Application Number
- CN202411162702.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-08-22
AI Technical Summary
Traditional infrared image detection methods for power transmission lines have low detection accuracy under complex backgrounds, making it difficult to effectively identify targets of varying sizes and uneven quantity, resulting in performance degradation in actual applications.
The improved YOLOv5s algorithm is adopted, and the global context attention mechanism and multiple context enhancement modules are introduced. Combined with the Wise-IoU v3 loss function, feature extraction and aggregation are optimized to improve the model's detection performance for small targets.
It improves the accuracy and robustness of infrared image detection in transmission lines, and can more accurately identify small targets such as insulators, overhanging wire clips, tension-resistant wire clips, etc., improving the reliability and accuracy of detection.
Smart Images

Figure CN119068308B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of infrared image recognition, and particularly to a multi-target detection method, device, equipment and medium for infrared images of transmission lines. Background Technique
[0002] Traditional inspection of transmission lines in power systems mainly relies on manual inspections. This method not only takes a long time, is costly, but also has safety problems. In recent years, using drones in combination with infrared thermal imagers to inspect transmission lines has gradually become a new inspection method. The infrared thermal imager can accurately capture the heat distribution of equipment, helping to detect areas with abnormal heat in advance to identify possible problems. Drones equipped with infrared thermal imagers can obtain thermal images of transmission lines in real time and accurately, thereby evaluating the operating conditions of transmission lines. Coupled with the continuous development of artificial intelligence, using drones in combination with deep learning algorithms to monitor power equipment on transmission lines, especially the excellent performance of the YOLO (You Only Look Once) algorithm in the field of target detection, makes it a powerful tool for processing images.
[0003] In the field of target detection, target detection algorithms based on deep learning are mainly divided into two categories. One category is single-stage target detection algorithms represented by the YOLO series algorithms and the SSD (Single Shot MultiBox Detector) algorithm, and the other category is two-stage target detection algorithms represented by the R-CNN (Region based Convolutional Neural Network) series algorithms. The advantage of the single-stage detection algorithms of the YOLO series lies in high real-time detection and fast detection speed. When processing the input image, it directly identifies and classifies and locates each region of the image. The two-stage target detection algorithm mainly means that the model divides the target detection task into two stages. The first stage is to generate candidate regions, and the second stage is to classify and regress the generated candidate regions. The advantage is higher detection accuracy, but slower speed. Therefore, the YOLO algorithm is mainly used for target detection in the current industrial inspection field.
[0004] To improve the detection accuracy of the YOLO series of algorithms, in the first prior art, by introducing a bidirectional feature pyramid network (BiFPN) to replace the original feature pyramid network in YOLOv5, the free exchange of information and the fusion of features are promoted, and the accuracy of the model is improved. In the second prior art, based on the YOLOv5s and YOLOv5x algorithms, an integrated model is constructed, and the weighted boxes fusion (WBF) inference technology is introduced, enabling the model to have good recognition effects on transmission towers under complex backgrounds, different lighting, and weather conditions, showing high robustness. The third prior art proposes an improved YOLOv5 algorithm for multi-object detection. Its main improvements include using ShuffleNetv2 as the main feature extraction structure to reduce the number of network parameters, and replacing the Bottleneck CSP module in the PANet network with the LightCSP module to accelerate feature fusion. In the fourth prior art, based on YOLOv5, lightweight Ghost convolution technology is adopted to replace the traditional convolution operation, and the original feature extraction network is replaced with a repeatedly weighted BiFPN network to achieve fast detection of insulator defects.
[0005] However, the above existing infrared image detection technologies for transmission lines all face some challenges: First, severe background interference or complex backgrounds lead to a decline in image quality, making it difficult for traditional detection methods to fully utilize image information, affecting the accurate expression of features by the model, and prone to false detections and missed detections. Second, when facing targets of different sizes and quantities, traditional models tend to identify easily classifiable samples with large sizes and large quantities, while ignoring difficult-to-classify samples with small sizes, long distances, and small quantities. As a result, the traditional model has a significant effect when simulating the algorithm on the PC side, but the performance of the model will be greatly reduced in actual applications, which urgently needs to be solved. Summary of the Invention
[0006] The present invention provides a multi-object detection method, device, equipment, and medium for infrared images of transmission lines to solve the problems of low detection accuracy of traditional detection methods and performance degradation in actual applications, and improve the accuracy and robustness of object detection.
[0007] In a first aspect embodiment of the present invention, a multi-object detection method for infrared images of transmission lines is provided. The method uses an improved YOLOv5s algorithm, which introduces a global context attention mechanism in the Backbone part of the original YOLOv5s algorithm and adds multiple context enhancement modules in the Neck part of the original YOLOv5s algorithm. Among them, the method includes the following steps:
[0008] Obtain an infrared image of a transmission line to be detected;
[0009] Based on the Backbone part of the improved YOLOv5s algorithm, calculate the attention weight feature information of the infrared image of the transmission line to be detected, and obtain the global feature relationship information according to the initial high-dimensional information and the attention weight feature information of the infrared image of the transmission line to be detected, and process the global feature relationship information, and obtain the enhanced global important information of the image according to the processed global feature relationship information and the initial high-dimensional information of the infrared image of the transmission line to be detected;
[0010] Based on the Neck part of the improved YOLOv5s algorithm, perform feature aggregation on the enhanced global important information of the image to obtain aggregated image information, and based on the Head part of the improved YOLOv5s algorithm, obtain the target detection result according to the aggregated image information.
[0011] According to an embodiment of the present invention, the target detection result includes at least one of the insulator position and insulator confidence, suspension clamp position and suspension clamp confidence, strain clamp position and strain clamp confidence, vibration damper position and vibration damper confidence, grading ring position and grading ring confidence, fault heating point position and fault heating point confidence.
[0012] According to an embodiment of the present invention, the Head part based on the improved YOLOv5s algorithm obtains the target detection result according to the aggregated image information, including:
[0013] Use the Head part of the improved YOLOv5s algorithm to generate detection frames, and perform classification, localization, and confidence scoring based on the detection frames and the aggregated image information to obtain the insulator position, insulator confidence, suspension clamp position, suspension clamp confidence, strain clamp position, strain clamp confidence, vibration damper position, vibration damper confidence, grading ring position, grading ring confidence, fault heating point position, and fault heating point confidence in the infrared image of the transmission line to be detected;
[0014] Obtain the target detection result according to the insulator position, insulator confidence, suspension clamp position, suspension clamp confidence, strain clamp position, strain clamp confidence, vibration damper position, vibration damper confidence, grading ring position, grading ring confidence, fault heating point position, and fault heating point confidence in the infrared image of the transmission line to be detected.
[0015] According to an embodiment of the present invention, the Neck part based on the improved YOLOv5s algorithm performs feature aggregation on the enhanced global important information of the image to obtain aggregated image information, including:
[0016] Based on the multiple context enhancement modules, perform adaptive operations on the globally important information of the enhanced image to obtain multiple adaptive weights;
[0017] Calculate the weighted sum of the multiple adaptive weights to obtain the aggregated image information.
[0018] According to an embodiment of the present invention, the improved YOLOv5s algorithm replaces the loss function of the original YOLOv5s algorithm with Wise-IoU v3. Before obtaining the infrared image of the transmission line to be detected, it further includes:
[0019] Obtain an infrared image set for training;
[0020] Perform labeling processing on the infrared image set to obtain a labeled data set. Among them, the labels are in YOLO format, and are divided into six labels: insulator, suspension clamp, strain clamp, vibration damper, grading ring, and fault heating point;
[0021] Based on a preset ratio, divide the labeled data set into a training set and a test set, use the training set to train the improved YOLOv5s algorithm, and after the training is completed, based on the Wise-IoU v3, use the test set to test the improved YOLOv5s algorithm to obtain a test result;
[0022] If the test result meets the preset test conditions, use the improved YOLOv5s algorithm to perform multi-object detection on the infrared image of the transmission line to be detected. Otherwise, adjust the preset ratio and retrain based on the adjusted training set until the new test result meets the preset test conditions.
[0023] According to the multi-object detection method for infrared images of transmission lines in the embodiments of the present invention, based on the Backbone part of the improved YOLOv5s algorithm, calculate the attention weight feature information of the infrared image of the transmission line to be detected, and obtain the global feature relationship information according to the initial high-dimensional information and the attention weight feature information, and obtain the globally important information of the enhanced image according to the processed global feature relationship information and the initial high-dimensional information of the infrared image of the transmission line to be detected; based on the Neck part of the improved YOLOv5s algorithm, perform feature aggregation on the globally important information of the enhanced image to obtain the aggregated image information, and based on the Head part of the improved YOLOv5s algorithm, obtain the target detection result according to the aggregated image information. Thereby, the problems of low detection accuracy and performance degradation in actual applications of traditional detection methods are solved, and the accuracy and robustness of target detection are improved.
[0024] In the second aspect of the present invention, an embodiment provides a multi-object detection device for infrared images of transmission lines. The device uses an improved YOLOv5s algorithm. The improved YOLOv5s algorithm introduces a global context attention mechanism in the Backbone part of the original YOLOv5s algorithm and adds multiple context enhancement modules in the Neck part of the original YOLOv5s algorithm. Among them, the device includes:
[0025] An acquisition module, configured to acquire an infrared image of a transmission line to be detected;
[0026] A processing module, configured to calculate attention weight feature information of the infrared image of the transmission line to be detected based on the Backbone part of the improved YOLOv5s algorithm, obtain global feature relationship information according to the initial high-dimensional information of the infrared image of the transmission line to be detected and the attention weight feature information, process the global feature relationship information, and obtain enhanced global important information of the image according to the processed global feature relationship information and the initial high-dimensional information of the infrared image of the transmission line to be detected;
[0027] A feature aggregation and detection module, configured to perform feature aggregation on the enhanced global important information of the image based on the Neck part of the improved YOLOv5s algorithm to obtain aggregated image information, and obtain a target detection result based on the aggregated image information according to the Head part of the improved YOLOv5s algorithm.
[0028] According to an embodiment of the present invention, the target detection result includes at least one of the insulator position and insulator confidence, suspension clamp position and suspension clamp confidence, strain clamp position and strain clamp confidence, vibration damper position and vibration damper confidence, grading ring position and grading ring confidence, and fault heating point position and fault heating point confidence.
[0029] According to an embodiment of the present invention, the feature aggregation and detection module is configured to:
[0030] Generate a detection box using the Head part of the improved YOLOv5s algorithm, and perform classification, positioning, and confidence scoring based on the detection box and the aggregated image information to obtain the insulator position, insulator confidence, suspension clamp position, suspension clamp confidence, strain clamp position, strain clamp confidence, vibration damper position, vibration damper confidence, grading ring position, grading ring confidence, fault heating point position, and fault heating point confidence in the infrared image of the transmission line to be detected;
[0031] The target detection result is obtained according to the insulator position, insulator confidence, suspension clamp position, suspension clamp confidence, strain clamp position, strain clamp confidence, vibration damper position, vibration damper confidence, grading ring position, grading ring confidence, fault heating point position and fault heating point confidence in the infrared image of the transmission line to be detected.
[0032] According to an embodiment of the present invention, the feature aggregation and detection module is used for:
[0033] Based on the multiple context enhancement modules, perform adaptive operations on the globally important information of the enhanced image to obtain multiple adaptive weights;
[0034] Calculate the weighted sum of the multiple adaptive weights to obtain the aggregated image information.
[0035] According to an embodiment of the present invention, the improved YOLOv5s algorithm replaces the loss function of the original YOLOv5s algorithm with Wise-IoU v3. Before obtaining the infrared image of the transmission line to be detected, the acquisition module is further used for:
[0036] Obtain an infrared image set for training;
[0037] Perform labeling processing on the infrared image set to obtain a labeled data set. Among them, the labels are in the YOLO format, and a total of six labels are divided, namely insulators, suspension clamps, strain clamps, vibration dampers, grading rings and fault heating points;
[0038] Based on a preset ratio, divide the labeled data set into a training set and a test set, use the training set to train the improved YOLOv5s algorithm, and after training is completed, based on the Wise-IoU v3, use the test set to test the improved YOLOv5s algorithm to obtain a test result;
[0039] If the test result meets the preset test conditions, use the improved YOLOv5s algorithm to perform multi-target detection on the infrared image of the transmission line to be detected. Otherwise, adjust the preset ratio and re-train based on the adjusted training set until the new test result meets the preset test conditions.
[0040] The multi-target detection device for infrared images of transmission lines according to an embodiment of the present invention is based on the Backbone part of the improved YOLOv5s algorithm, calculates the attention weight feature information of the infrared image of the transmission line to be detected, obtains the global feature relationship information according to the initial high-dimensional information and the attention weight feature information, and obtains the enhanced global important information of the image according to the processed global feature relationship information and the initial high-dimensional information of the infrared image of the transmission line to be detected; based on the Neck part of the improved YOLOv5s algorithm, performs feature aggregation on the enhanced global important information of the image to obtain aggregated image information, and based on the Head part of the improved YOLOv5s algorithm, obtains the target detection result according to the aggregated image information. Thereby, the problems of low detection accuracy of traditional detection methods and performance degradation in practical applications are solved, and the accuracy and robustness of target detection are improved.
[0041] The third aspect of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the multi-target detection method for infrared images of transmission lines as described in the above embodiments.
[0042] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the multi-target detection method for infrared images of transmission lines as described in the above embodiments.
[0043] The additional aspects and advantages of the present invention will be partly given in the following description, partly become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0044] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0045] Figure 1 is the structure diagram of the YOLOv5s network;
[0046] Figure 2 is the schematic diagram of the GC block network structure;
[0047] Figure 3 is the comparison diagram of the C3GC and C3 network structures;
[0048] Figure 4 is the structure diagram of the CAM (Context Enhancement) module according to an embodiment of the present invention;
[0049] Figure 5 is the structure diagram of the improved YOLOv5s algorithm according to an embodiment of the present invention;
[0050] Figure 6 Flow chart of a multi-target detection method for infrared images of transmission lines provided according to an embodiment of the present invention;
[0051] Figure 7 Schematic diagram of Wise-IoU regression;
[0052] Figure 8 Schematic diagram of an infrared grayscale image according to an embodiment of the present invention;
[0053] Figure 9 Schematic diagram of the CAM fusion mechanism according to an embodiment of the present invention;
[0054] Figure 10 Comparison chart of ablation experiment mAP@0.5 according to an embodiment of the present invention;
[0055] Figure 11 Effect diagram of the loss function according to an embodiment of the present invention;
[0056] Figure 12 Schematic diagram of an infrared image sample according to an embodiment of the present invention;
[0057] Figure 13 Effect diagram of the prediction of the original YOLOv5s algorithm according to an embodiment of the present invention;
[0058] Figure 14 Effect diagram of the prediction of the improved YOLOv5s algorithm according to an embodiment of the present invention;
[0059] Figure 15 Block diagram of a multi-target detection device for infrared images of transmission lines according to an embodiment of the present invention;
[0060] Figure 16 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention.
[0061] Reference numerals: 10 - multi-target detection device for infrared images of transmission lines, 100 - acquisition module, 200 - processing module, 300 - feature aggregation and detection module; 1601 - memory, 1602 - processor, 1603 - communication interface. Detailed implementation manners
[0062] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.
[0063] The multi-object detection method, device, equipment and medium of the infrared image of the transmission line according to the embodiment of the present invention will be described below with reference to the accompanying drawings.
[0064] Before introducing the multi-object detection method of the infrared image of the transmission line according to the embodiment of the present invention, the YOLOv5 algorithm, the YOLOv5s network structure, and the improved YOLOv5s algorithm involved in the multi-object detection method of the infrared image of the transmission line proposed by the present invention will be introduced first.
[0065] First, introduce the YOLOv5 algorithm. The YOLOv5 algorithm is a relatively popular single-stage object detection algorithm at the present stage. It has made significant improvements in both speed and lightweight design. Its model is small, and it is more practical than algorithms such as YOLOv7 and YOLOv8 that appeared later. Compared with previous algorithms such as YOLOv4 and YOLOv3, it improves the detection accuracy and model performance while maintaining high speed. Among the YOLOv5 series of algorithms, the YOLOv5s model is the smallest and has the fastest detection speed, making it suitable for handling tasks with high real-time requirements.
[0066] Secondly, introduce the YOLOv5s network structure. The network structure of YOLOv5s is as shown in Figure 1 (a). It mainly consists of three parts: the Backbone part, the Neck part, and the Head part. The Backbone part mainly consists of the following key parts: the convolutional layer (Conv), the C3 module, and the Spatial Pyramid Pooling Fast (SPPF) module. The convolutional layer combines the batch normalization layer and the Silu activation function to form the Conv module, which is used to extract target features. The network structure of the C3 module is as shown in Figure 1 (b). The C3 module enhances the feature learning ability and alleviates the problem of gradient disappearance in the deep network by combining the Bottleneck structure for residual learning and the Concat operation. The network structure of the Bottleneck module in it is as shown in Figure 1 (c). Its main principle is to combine 1x1 and 3x3 convolutional layers to achieve dimensionality reduction and dimensionality increase of the feature channels, thereby reducing the computational complexity while maintaining the feature expression ability. The SPPF module is as shown in Figure 1 (d). It improves the processing speed and optimizes the feature capture of multi-scale objects by using multiple small-size pooling kernels. The Neck part adopts the Feature Pyramid Networks (FPN) structure and is supplemented by the PANet structure to strengthen feature fusion and location information transmission. The Head part is responsible for generating detection boxes and performing classification, location, and confidence scoring on them.
[0067] Finally, the improved YOLOv5s algorithm proposed by the present invention is introduced. Among them, the global context attention mechanism is introduced in the Backbone part of the original YOLOv5s algorithm, and multiple context enhancement modules are added to the Neck part of the original YOLOv5s algorithm.
[0068] Specifically, the introduction of the global context attention mechanism in the Backbone part of the original YOLOv5s algorithm is described first.
[0069] The global context attention mechanism (Global Context Block, GC) is derived from the GC Net paper. The GC block is improved on the basis of the non-local network. Experiments have shown that capturing long-range dependencies is beneficial to various recognition tasks. In traditional neural networks, long-range dependencies can only be modeled by continuously stacking convolutional layers, which is inefficient and computationally intensive. For this reason, the GC block with better performance is proposed.
[0070] Figure 2 For the network structure of the GC block, the input image to be detected first calculates the attention weight feature information of the image through a 1×1 convolutional block and a Softmax module, and then calculates the global feature relationship information of C×1×1 with the input H×W×C, as shown in Equation (1); then, two 1×1 convolutional blocks in the Transform structure are used to reduce the model parameters and save the computational overhead. Moreover, a LayerNorm module is inserted between the two 1×1 convolutional blocks, which can improve the training stability of the model and has a certain regularization effect, reducing the risk of model overfitting; finally, the initial information H×W×C is added to the global feature relationship information C×1×1 to obtain the output result of the enhanced global important information of the image, as shown in Equation (3).
[0071]
[0072] Among them, α j represents the weight of global attention pooling, represents the attention score in the calculation of the attention mechanism, W k represents the weight matrix, x j represents the input vector containing the feature information of a certain position, N p is the total number of features, x m represents the input vector containing the global features after pooling or aggregation, m represents the dimension size (number of features) of the input feature vector, represents the input vector x m after passing through the weight matrix W k weighted attention score, z krepresents the global context attention mechanism modeling architecture, F(·,·) represents the fusion function that aggregates global context features into features for each position, and x k represents the input vector containing the intermediate feature results obtained in computing the global attention, δ(·) represents the feature transformation that captures channel dependencies, and x i represents the input vector containing the features of a specific position, ReLU() represents the activation function, LN() represents a normalization method that normalizes the input of each layer, and W v1 represents the weight matrix, which performs a preliminary linear transformation on the features, and W v2 represents the weight matrix, which represents a further linear transformation on the features after normalization and the activation function, and z i represents the input vector x i and the total feature vector obtained after fusing with the global context feature information obtained after transformation.
[0073] Equation (1) represents the weight of global attention pooling, Equation (2) represents the global context attention mechanism modeling architecture, and Equation (3) is the specific calculation formula of GC. Specifically, the calculation of GC mainly includes three parts. The first step is to calculate the weight α of global attention pooling j ; the second step is to capture channel dependencies through the bottleneck transformation δ(·); the third step is to perform feature fusion using the fusion function F(·,·).
[0074] Furthermore, Figure 3 a detailed comparison is made between the new module C3GC proposed in the present invention and the existing C3 module. In the Backbone part, as the number of network layers increases, the C3 module faces challenges in extracting shallow features and is prone to the problem of gradient disappearance when the network depth increases. In view of this, the present invention cleverly embeds a GC block before one of the Conv modules, and the other Conv module remains unchanged. Compared with embedding GC blocks before both Conv modules, this improved method proposed in the present invention can introduce global context information according to specific requirements while maintaining the simplicity and effectiveness of model calculation. Choosing to fuse the GC block in one branch, the branch with the other GC block can continue to focus on learning fine-grained local features, thus forming an effective complement between context information and detail information, and can achieve an organic combination of global context and local features, thereby improving the performance of the model in object detection and image processing tasks.
[0075] It is understandable that during the detection of transmission lines, there are some small target devices, and the captured images are restricted by the environment, shooting distance, shooting angle, etc., resulting in uneven infrared image quality. Therefore, in the Neck network part of the original YOLOv5s algorithm, this invention introduces the CAM (Context Augmentation Module) module to balance the training process, reduce feature conflicts, improve the detection accuracy of small targets, and enhance the robustness of the model while maintaining computational efficiency. Without changing the scale of the feature map, the CAM module expands the scale of the shared convolutional kernel to increase the receptive field of the convolutional kernel.
[0076] The CAM module applies dilated convolutions with different dilation rates to obtain context information with different receptive fields and injects it into the FPN (Feature Pyramid Networks) network in YOLOv5 from top to bottom. Its structural diagram is as Figure 4 shown. The features are processed by dilated convolution with a 3×3 convolutional kernel at different dilation ratios of 1, 3, and 5, so that the module realizes the purpose of multi-scale extraction of the feature map.
[0077] Thus, by introducing the global context attention mechanism in the Backbone part of the original YOLOv5s algorithm and introducing the context augmentation module in the Neck part of the original YOLOv5s algorithm, this invention proposes an improved YOLOv5s algorithm as Figure 5 shown.
[0078] As Figure 5 shown, Figure 5 the part circled by the dashed box in the figure is the module added by this invention on the basis of the original YOLOv5s. First, in order to enhance the model's ability to extract global features and reduce computational overhead, this invention embeds 4 C3GC modules in the Backbone part of the original YOLOv5s to extract features and improve the ability to extract global context features. Second, three CAM modules are added before the three concat modules in the Neck part of the original YOLOv5s. The CAM module is used to enhance the detection performance of small targets. By applying dilated convolutions with different dilation rates to obtain context information with different receptive fields, it realizes multi-scale extraction of the feature map, and through three mechanisms of weighted fusion, cascading operation, and adaptive operation, aggregates the context information into the output feature map, thereby reducing information conflicts and improving the fusion effect of the feature map.
[0079] Therefore, the improved YOLOv5s algorithm proposed by the present invention first introduces a global context attention mechanism in the Backbone part of the original YOLOv5s to enhance the model's ability to extract global features and reduce its computational overhead. Secondly, a context enhancement module is introduced in the Neck part of the original YOLOv5s to improve the detection performance for small targets. The improved YOLOv5s algorithm proposed by the present invention solves the problems of low-quality data sets, dense and various-sized devices in the operation detection of transmission lines.
[0080] The following introduces the multi-object detection method for infrared images of transmission lines using the above improved YOLOv5s algorithm proposed by the present invention.
[0081] Specifically, Figure 6 is a schematic flow chart of a multi-object detection method for infrared images of transmission lines provided by an embodiment of the present invention.
[0082] As Figure 6 shown, the multi-object detection method for infrared images of transmission lines includes the following steps:
[0083] In step 601, an infrared image of a transmission line to be detected is obtained.
[0084] Among them, the infrared image of the transmission line to be detected refers to image data reflecting the surface temperature distribution of the transmission line to be detected and its attached equipment.
[0085] Specifically, in the embodiment of the present invention, an infrared thermal imager in the prior art can be used to obtain the infrared image of the transmission line to be detected. In addition, those skilled in the art can also use other methods in the prior art to obtain the infrared image of the transmission line to be detected according to the actual situation, which is not specifically limited herein.
[0086] Further, in some embodiments, the improved YOLOv5s algorithm replaces the loss function of the original YOLOv5s algorithm with Wise-IoU v3. Before obtaining the infrared image of the transmission line to be detected, it further includes: obtaining an infrared image set for training; performing labeling processing on the infrared image set to obtain a labeled data set, where the labels are in YOLO format and are divided into six labels: insulator, suspension clamp, strain clamp, vibration damper, grading ring, and fault heating point; based on a preset ratio, dividing the labeled data set into a training set and a test set, training the improved YOLOv5s algorithm using the training set, and testing the improved YOLOv5s algorithm using the test set based on Wise-IoU v3 after training to obtain a test result; if the test result meets the preset test conditions, using the improved YOLOv5s algorithm to perform multi-object detection on the infrared image of the transmission line to be detected, otherwise, adjusting the preset ratio and retraining based on the adjusted training set until the new test result meets the preset test conditions.
[0087] It can be understood that object detection is crucial in computer vision, and its accuracy depends on the design of the loss function. The bounding box loss function is crucial for improving the model performance. However, the object detection methods for infrared images of transmission lines are usually trained and optimized based on high-quality data sets. In practice, environmental factors may cause the training data to contain a large number of ordinary and low-quality samples. If these samples are over-optimized, it will damage the model performance. Therefore, the embodiment of the present invention replaces the loss function of YOLOv5s with Wise-IoU, which introduces a dynamic non-linear focusing mechanism, evaluates the anchor boxes through "outlier degree", intelligently allocates gradient gains, reduces the dependence on low-quality samples, and improves the detector performance. There are three versions of Wise-IoU, namely v1, v2, and v3 versions. Wise-IoU v1 adopts a two-level attention mechanism based on distance, and its calculation formula is:
[0088]
[0089] Where, is the bounding box loss under the Wise-IoU v1 loss function, is the bounding box loss under IoU, R WloU is the distance attention, (x, y) is the predicted box of the target, (x gt , y gt ) is the ground truth box of the target, W g and H g are the width and height of the minimum bounding rectangle of the predicted box and the ground truth box respectively. The superscript * indicates separating W g and H g from the computational graph, which can effectively improve the convergence efficiency. The regression schematic diagram is shown in Figure 7 .
[0090] Compared with Wise-IoU v1, Wise-IoU v3 adopted in the embodiments of the present invention does not have calculations related to aspect ratio and distance attention. Instead, it adopts a dynamic non-monotonic focusing mechanism and introduces "outlier degree" to describe the quality of anchor boxes. When the outlier degree of an anchor box is small, it indicates that its quality is high, and the model will assign a small gradient gain to it. At the same time, when the outlier degree of the anchor box is extremely large (the quality of the anchor box is extremely low), the model also assigns a small gradient gain to it to prevent it from generating large harmful gradients, so that the model can focus more on ordinary and lower-quality anchor boxes, improving the detection performance of the model in complex backgrounds. The calculation formula of Wise-IoU v3 is as follows:
[0091]
[0092] Among them, r in formula (6) is the non-monotonic focusing parameter, β is the outlier degree, and δ and α are hyperparameters. In the present invention, the hyperparameters δ and α are set to 3 and 1.9 respectively. is the moving average, is the gradient gain, and in formula (8) is the bounding box loss under the Wise-IoU v3 loss function. And, since is updated in real time, this enables Wise-IoU v3 to dynamically adjust the gradient gain according to the current training status.
[0093] Thus, by replacing the loss function with Wise-IoU v3 and adopting a dynamic non-monotonic focusing mechanism, the model focuses more on ordinary and lower-quality anchor boxes, improving the detection performance of the model and making the model more practical.
[0094] In summary, the improved YOLOv5s algorithm in the embodiments of the present invention uses Wise-IoU v3 to replace the loss function in YOLOv5s because Wise-IoU v3 not only considers the overlapping area of the target box, but also focuses on anchor boxes of ordinary quality. Therefore, it can better handle the situation where the target boundary is inaccurate. This method can effectively improve the target detection accuracy of the model in complex backgrounds, especially perform well in scenarios with many small targets and occlusions, improving the positioning performance of multiple devices in the transmission line of the present invention.
[0095] Furthermore, the improved YOLOv5s algorithm is trained and tested, and when the test results meet the preset test conditions, the improved YOLOv5s algorithm is used to perform multi-target detection on the infrared image of the transmission line to be detected.
[0096] Specifically, first, an infrared image set for training is obtained using infrared thermal imaging technology, and the infrared images in the infrared image set are grayscale processed. The infrared images after grayscale processing are as Figure 8 shown. The grayscale processed infrared image set is labeled using a labeling tool (such as the LabelImg tool) to obtain a labeled data set. The labels are in the YOLO format, and a total of six labels are divided: insulator, suspension clamp, strain clamp, shockproof hammer, grading ring, and fault heating point.
[0097] Furthermore, the labeled data set is divided into a training set and a test set according to a preset ratio. For example, taking the number of infrared images in the infrared image set as 1270 and the preset ratio as 7:3, the number of the training set, test set, and various labels after division is shown in Table 1.
[0098] Table 1
[0099] Name Quantity Training set 889 Test set 381 insulator 1845 X clamp 482 N clamp 249 hammer 1168 ring 280 error 253
[0100] Among them, insulator in Table 1 represents the insulator in the infrared image, X clamp represents the suspension clamp in the infrared image, N clamp represents the strain clamp in the infrared image, hammer represents the shockproof hammer in the infrared image, ring represents the grading ring in the infrared image, and error represents the possible fault heating point in the infrared image.
[0101] Furthermore, the improved YOLOv5s algorithm is used as the basic model, and the model is trained using the training set. The training environment of the embodiment of the present invention is under the Windows 10 operating system, using NVIDIA GeForce GTX 1650 as the hardware, Pytorch as the deep learning framework, Python version 3.8, CUDA version 12.1, and the training epoch is set to 100. Among them, the momentum is 0.937, the weight decay is 0.0005, the batch size is 8, the learning rate is 0.0001, and the input image size is 640×640.
[0102] Furthermore, after the model training is completed, each image in the test set is traversed. For each image, its true label is obtained, and for each prediction box in each image, the intersection over union (IoU) between it and the true box is calculated using Wise-IoU v3. If the IoU exceeds a certain threshold (such as 0.8), then the prediction is determined to be a correct match.
[0103] Further, count the number of correct matches for each category, and use evaluation metrics such as Precision, Recall, and Mean Average Precision (mAP) to evaluate the performance of the model on the test set.
[0104] Among them, Precision refers to the proportion of positive samples correctly predicted by the model. The calculation formula is shown in Equation (9).
[0105]
[0106] Among them, P is Precision, T P is the number of positive samples correctly predicted, and F P is the number of positive samples wrongly predicted.
[0107] Recall refers to the proportion of positive samples predicted by the model among all positive samples. The calculation formula is shown in Equation (10), where F N represents the number of negative samples wrongly predicted by the model.
[0108]
[0109] Among them, R is Recall, T P is the number of positive samples correctly predicted, and F N is the number of negative samples wrongly predicted by the model.
[0110] Mean Average Precision is obtained by averaging the Average Precision (AP) of a single category. The calculation formula is shown in Equation (11). AP is determined by calculating the area under the P-R curve, and the calculation formula is shown in Equation (12).
[0111]
[0112] Among them, L mAP is Mean Average Precision, L AP is Average Precision, n is the number of segments on the P-R curve, that is, the number of data points for calculating Mean Average Precision, and P(R) is the Precision-Recall (P-R) curve function.
[0113] Furthermore, the test results are obtained based on precision, recall, and mean average precision. If the test results meet the preset test conditions, such as when the precision, recall, and mean average precision reach specific thresholds, it is determined that the test results meet the preset test conditions, that is, the model training is successful, and the improved YOLOv5s algorithm is used to perform multi-object detection on the infrared image of the transmission line to be detected. If the test results do not meet the preset test conditions, it means that the performance of the model does not meet the expectations, and the division ratio of the dataset needs to be adjusted, and retraining is performed based on the adjusted training set until the new test results meet the preset test conditions.
[0114] It should be noted that in each round of training in the embodiments of the present invention, the loss result is calculated using Wise-IoU v3, and the loss result is fed back to the improved YOLOv5s model to update the parameters of the loss function, and the next round of training is performed after the update of the parameters of the loss function is completed.
[0115] In step S602, based on the Backbone part of the improved YOLOv5s algorithm, the attention weight feature information of the infrared image of the transmission line to be detected is calculated, and the global feature relationship information is obtained according to the initial high-dimensional information and the attention weight feature information of the infrared image of the transmission line to be detected, and the global feature relationship information is processed, and the enhanced global important information of the image is obtained according to the processed global feature relationship information and the initial high-dimensional information of the infrared image of the transmission line to be detected.
[0116] Specifically, the infrared image of the transmission line to be detected is input into the Backbone part of the improved YOLOv5s algorithm, and through Figure 2 the 1×1 convolutional block and the Softmax module in the GC block shown in the figure, the attention weight feature information of the infrared image of the transmission line to be detected is calculated, and according to the initial high-dimensional information H×W×C of the input infrared image of the transmission line to be detected, the global feature relationship information of C×1×1 is calculated.
[0117] Furthermore, as Figure 2 shown in the figure, the number of model parameters is reduced by two 1×1 convolutional blocks in the Transform structure, saving computational overhead, and a LayerNorm module is inserted between the two 1×1 convolutional blocks, which can improve the training stability of the model and has a certain regularization effect, reducing the risk of model overfitting. Finally, the initial high-dimensional information H×W×C of the infrared image of the transmission line to be detected is added to the global feature relationship information C×1×1 to obtain the output result of the enhanced global important information of the image.
[0118] In step S603, based on the Neck part of the improved YOLOv5s algorithm, feature aggregation is performed on the globally important information of the enhanced image to obtain aggregated image information, and based on the Head part of the improved YOLOv5s algorithm, the target detection result is obtained according to the aggregated image information.
[0119] Among them, in some embodiments, the target detection result includes at least one of the insulator position and insulator confidence, suspension clamp position and suspension clamp confidence, strain clamp position and strain clamp confidence, vibration damper position and vibration damper confidence, grading ring position and grading ring confidence, and fault heating point position and fault heating point confidence.
[0120] Further, in some embodiments, based on the Neck part of the improved YOLOv5s algorithm, feature aggregation is performed on the globally important information of the enhanced image to obtain aggregated image information, including: performing adaptive operations on the globally important information of the enhanced image based on multiple context enhancement modules to obtain multiple adaptive weights; calculating the weighted sum of the multiple adaptive weights to obtain the aggregated image information.
[0121] Among them, the context enhancement module of the embodiment of the present invention includes 3 fusion mechanisms, as Figure 9 shown, which are weighted fusion, cascade operation, and adaptive operation, Figure 9 (a) and Figure 9 (b) The fusion mechanisms represented are both to adjust the number of channels of the feature map through 3 1×1 convolutions, and then perform feature fusion in the spatial and channel dimensions, while Figure 9 (c) The adaptive operation represented is to obtain the adaptive weight through convolution, splicing, and the Softmax function, and through calculating the weighted sum, the context information can be aggregated to the output. The connection method of the context enhancement module of the embodiment of the present invention is Figure 9 the Adaptive connection method represented by (c).
[0122] Specifically, input the globally important information of the enhanced image into the Neck part of the improved YOLOv5s algorithm, adjust the number of channels of the enhanced image through a convolutional layer (such as 3 1×1 convolutions), and perform feature fusion in the spatial and channel dimensions. Perform adaptive operations on the globally important information of the enhanced image through convolution, splicing, and the Softmax function. For each context enhancement module, perform weighted summation on the enhanced image according to the output adaptive weight to obtain the aggregated output of the context enhancement module. Perform weighted summation on the aggregated results output by all context enhancement modules to obtain the final aggregated image information.
[0123] Furthermore, in some embodiments, based on the Head part of the improved YOLOv5s algorithm, the target detection result is obtained according to the aggregated image information, including: generating detection boxes using the Head part of the improved YOLOv5s algorithm, and performing classification, localization, and confidence scoring based on the detection boxes and the aggregated image information to obtain the positions of insulators, insulator confidence levels, positions of suspension clamps, suspension clamp confidence levels, positions of strain clamps, strain clamp confidence levels, positions of dampers, damper confidence levels, positions of grading rings, grading ring confidence levels, positions of fault heating points, and fault heating point confidence levels in the infrared image of the transmission line to be detected; obtaining the target detection result according to the positions of insulators, insulator confidence levels, positions of suspension clamps, suspension clamp confidence levels, positions of strain clamps, strain clamp confidence levels, positions of dampers, damper confidence levels, positions of grading rings, grading ring confidence levels, positions of fault heating points, and fault heating point confidence levels in the infrared image of the transmission line to be detected.
[0124] Specifically, the aggregated image information is input into the Head part of the improved YOLOv5s algorithm, and the Head part of the improved YOLOv5s algorithm is used to generate detection boxes, and based on the detection boxes and the aggregated image information, the category (such as insulator, suspension clamp, strain clamp, etc.) to which each detection box belongs, the center coordinates, width-to-height ratio of each detection box, and the confidence score that each detection box belongs to a certain category are predicted.
[0125] Furthermore, for each detected insulator, its position (coordinates) and confidence score in the infrared image of the transmission line to be detected are output. Similarly, for the detected positions of suspension clamps, strain clamps, dampers, grading rings, and fault heating points, the positions of suspension clamps, suspension clamp confidence levels, positions of strain clamps, strain clamp confidence levels, positions of dampers, damper confidence levels, positions of grading rings, grading ring confidence levels, positions of fault heating points, and fault heating point confidence levels are output respectively.
[0126] Finally, the target detection result is obtained according to the positions of insulators, insulator confidence levels, positions of suspension clamps, suspension clamp confidence levels, positions of strain clamps, strain clamp confidence levels, positions of dampers, damper confidence levels, positions of grading rings, grading ring confidence levels, positions of fault heating points, and fault heating point confidence levels in the infrared image of the transmission line to be detected.
[0127] In addition, in order to verify the effectiveness of the improvements such as adding the C3GC module, adding the CAM module, and replacing the Wise-IoU loss function to the original YOLOv5s algorithm, ablation experiments were conducted and the results were compared. The number of rounds of the ablation experiment and mAP@0.5 are as Figure 10 shown, Figure 10 where the abscissa represents the number of training rounds and the ordinate represents mAP@0.5. In Figure 10Among them, the blue curve YOLOv5s represents the original YOLOv5s algorithm, the orange curve YOLOv5s+Wise-iou represents the improved algorithm after replacing with Wise-IoU v3, the green curve YOLOv5s+CAM represents the improved algorithm after adding the CAM module, the red curve YOLOv5s+C3GC represents the improved algorithm after adding the C3GC module, and the purple curve YOLOv5s+ECW represents the improved YOLOv5s algorithm after integrating the above three improvement methods.
[0128] As can be seen from Figure 10 it, compared with the original YOLOv5s algorithm, the improved YOLOv5s algorithm not only improves in accuracy, but also the convergence speed of the curve is faster than that of the original YOLOv5s curve. The curve of YOLOv5s+ECW (the improved YOLOv5s algorithm) is smoother after reaching the accuracy peak.
[0129] Specific data are shown in Table 2. The mAP@0.5 / % of Model 1, Model 2 and Model 3 increases by 4.0%, 2.6%, 1.8% respectively compared with the original YOLOv5s algorithm. Model 4, the final improved YOLOv5s algorithm after integrating the three improvement points, is 4.8% higher than the original YOLOv5s algorithm in mAP@0.5 / %, and indicators such as precision and recall have been improved to varying degrees, proving the effectiveness and feasibility of the improved YOLOv5s algorithm.
[0130] Table 2
[0131] Model YOLOv5s +C3GC +CAM +Wise-IoU P / % R / % mAP@0.5 / % 0 √ 76.5 67.0 69.3 1 √ √ 77.6 71.3 72.6 2 √ √ 80.7 68.1 71.2 3 √ √ 76.0 68.7 70.4 4 √ √ √ √ 84.5 71.5 73.2
[0132] Furthermore, Figure 11 it is a comparison chart of the loss function effects of the original YOLOv5s algorithm and the improved YOLOv5s algorithm proposed by the present invention. Figure 11 (a) is the training classification loss function image. Figure 11 (b) is the validation classification loss function image. Figure 11 In it, the orange curve is the loss function of the original YOLOv5s algorithm, and the blue curve is the loss function of the improved YOLOv5s algorithm. As can be seen from Figure 11 it, the improved YOLOv5s algorithm has a smaller loss value compared with the original YOLOv5s, and the curve tends to be stable faster, especially the effect on the test set is more obvious, reflecting the effectiveness of replacing the Wise-IoU v3 loss function in the present invention.
[0133] Furthermore, the present invention selects four images from infrared images to test the detection effects of the original YOLOv5s algorithm and the improved YOLOv5s algorithm. As shown in Figure 12 , Figure 13 and Figure 14 shown.Figure 12 is an infrared image sample diagram, Figure 12 including four infrared image sample diagrams (a), (b), (c), and (d), Figure 13 is the effect diagram after predicting (a), Figure 12 (a), Figure 12 (b), Figure 12 (c), and Figure 12 (d) using the YOLOv5s algorithm, Figure 14 is the effect diagram after predicting (a), Figure 12 (a), Figure 12 (b), Figure 12 (c), and Figure 12 (d) using the improved YOLOv5s algorithm. Figure 13 And Figure 14 display the component category and confidence level on the rectangular frame for component positioning. insulator represents insulator, X clamp represents suspension clamp, N clamp represents strain clamp, hammer represents vibration damper, ring represents grading ring, and error represents possible fault heating points in the infrared image.
[0134] When the original YOLOv5s algorithm detects images, due to the relatively dark background and the small size of the equipment, some components of power equipment are not detected. Moreover, for the detection of small targets such as vibration dampers, strain clamps, and suspension clamps by the original YOLOv5s algorithm, the confidence level of the rectangular frame is not high. This is because before the original YOLOv5s was improved, when performing multi-target detection of transmission lines, it tended to identify targets with a large number and large size such as insulators, ignoring the attention to small targets such as vibration dampers and strain clamps, resulting in low detection accuracy for small targets. However, the improved YOLOv5s algorithm proposed in the present invention avoids this problem. Comparing Figure 12 and Figure 13 it can be seen that Figure 12 in the leftmost image (a) of the four images, the original YOLOv5s algorithm missed a vibration damper during detection, Figure 12 in (b), the original YOLOv5s algorithm identified two vibration dampers and two strain clamps as one. Therefore, the problem of missing detection of small targets by the original YOLOv5s algorithm is serious.
[0135] The improved YOLOv5s algorithm proposed by the present invention focuses more on small targets, effectively avoiding the occurrence of missed detection problems. The detection of small targets such as vibration dampers and strain clamps is more accurate and has a higher confidence level compared to the original YOLOv5s algorithm. Moreover, due to the adoption of the context attention module and the use of the multi-scale fusion mechanism in the method proposed by the present invention, the extraction and aggregation of features are significantly superior to the original YOLOv5s algorithm. In other words, the improved YOLOv5s algorithm performs better on low-quality datasets. In practice, the quality of infrared images captured varies due to environmental, brightness, and distance limitations, so there is a greater need for the algorithm's ability to extract and aggregate features, which intuitively demonstrates the effectiveness of the method proposed by the present invention in the field of transmission line image detection.
[0136] Furthermore, Table 3 shows the recognition of various targets by the original YOLOv5s algorithm and the improved YOLOv5s algorithm of the present invention. The improved YOLOv5s algorithm proposed by the present invention does not significantly improve the recognition effect of insulators. However, for small targets that are not easily detected, such as strain clamps and suspension clamps, the detection accuracy of the improved YOLOv5s algorithm for strain clamps, suspension clamps, and vibration dampers has increased by 5.6%, 7.9%, and 3.3% respectively compared to the original YOLOv5s algorithm, further verifying the effectiveness of the improved YOLOv5s algorithm.
[0137] Table 3
[0138]
[0139] Furthermore, to verify the superiority of the improved YOLOv5s algorithm of the present invention, it was compared with various mainstream detection algorithms for the dataset of the present invention, as shown in Table 4. It can be seen that the performance of the original YOLOv5s algorithm has been improved to varying degrees compared to the SSD and YOLOv7 algorithms. However, when compared with algorithms such as Faster R-CNN and YOLOv8, although the accuracies of the two algorithms are similar to that of the original YOLOv5s algorithm, the original YOLOv5s algorithm is less than Faster R-CNN and YOLOv8. Therefore, the original YOLOv5s algorithm requires fewer computing resources for deployment and operation. Therefore, overall comparison shows that the original YOLOv5s algorithm is more suitable for industrial practical applications. Therefore, the present invention adopts the original YOLOv5s algorithm and improves it on this basis. Data shows that the improved YOLOv5s algorithm (YOLOv5s-ECW in Table 4) has improved in indicators such as precision and recall compared to before, and can well meet the requirements for infrared image detection of transmission lines.
[0140] Table 4
[0141]
[0142]
[0143] Therefore, the present invention proposes an improved YOLOv5s algorithm based on the original YOLOv5s for multi-object detection of infrared images of transmission lines. First, a C3GC module is introduced in the Backbone part to enhance the model's ability to extract global context features. Secondly, a CAM context enhancement module is introduced in the Neck part, and finally, a new loss function, Wise-IoU, is replaced. The test results show that compared with the original YOLOv5s algorithm, the improved YOLOv5s algorithm has an average precision improvement of 4.8%, and at the same time, the precision, recall rate, etc. have all been improved to varying degrees. Moreover, the improved YOLOv5s algorithm can focus more on the detection of small targets and better meet the practical requirements than the original YOLOv5s.
[0144] According to the multi-object detection method for infrared images of transmission lines in an embodiment of the present invention, based on the Backbone part of the improved YOLOv5s algorithm, the attention weight feature information of the infrared image of the transmission line to be detected is calculated, and the global feature relationship information is obtained according to the initial high-dimensional information and the attention weight feature information. And according to the processed global feature relationship information and the initial high-dimensional information of the infrared image of the transmission line to be detected, the enhanced global important information of the image is obtained; based on the Neck part of the improved YOLOv5s algorithm, the enhanced global important information of the image is subjected to feature aggregation to obtain aggregated image information, and based on the Head part of the improved YOLOv5s algorithm, the target detection result is obtained according to the aggregated image information. Thus, the problem that in the infrared image detection of transmission lines, due to factors such as the environment, shooting angle, and shooting distance, the quality of the infrared image dataset is not high, and the number of recognized targets is unbalanced and the target sizes are different, resulting in low recognition accuracy is solved, and the reliability and accuracy of the infrared image detection of transmission lines are improved.
[0145] Secondly, a multi-object detection device for infrared images of transmission lines proposed according to an embodiment of the present invention is described with reference to the accompanying drawings.
[0146] Figure 15 It is a block diagram of a multi-object detection device for infrared images of transmission lines in an embodiment of the present invention.
[0147] In this embodiment, the device uses the improved YOLOv5s algorithm. The improved YOLOv5s algorithm introduces a global context attention mechanism in the Backbone part of the original YOLOv5s algorithm and adds multiple context enhancement modules in the Neck part of the original YOLOv5s algorithm.
[0148] Such as Figure 15As shown in the figure, the multi-object detection device 10 for the infrared image of the transmission line includes: an acquisition module 100, a processing module 200, and a feature aggregation and detection module 300.
[0149] Among them, the acquisition module 100 is used to acquire the infrared image of the transmission line to be detected; the processing module 200 is used to calculate the attention weight feature information of the infrared image of the transmission line to be detected based on the Backbone part of the improved YOLOv5s algorithm, and obtain the global feature relationship information according to the initial high-dimensional information and the attention weight feature information of the infrared image of the transmission line to be detected, and process the global feature relationship information, and obtain the enhanced global important information of the image according to the processed global feature relationship information and the initial high-dimensional information of the infrared image of the transmission line to be detected; the feature aggregation and detection module 300 is used to perform feature aggregation on the enhanced global important information of the image based on the Neck part of the improved YOLOv5s algorithm to obtain the aggregated image information, and based on the Head part of the improved YOLOv5s algorithm, obtain the target detection result according to the aggregated image information.
[0150] Further, in some embodiments, the target detection result includes at least one of the insulator position and insulator confidence, suspension clamp position and suspension clamp confidence, strain clamp position and strain clamp confidence, vibration damper position and vibration damper confidence, grading ring position and grading ring confidence, fault heating point position and fault heating point confidence.
[0151] Further, in some embodiments, the feature aggregation and detection module 300 is used to: generate detection frames using the Head part of the improved YOLOv5s algorithm, and perform classification, localization, and confidence scoring based on the detection frames and the aggregated image information to obtain the insulator position, insulator confidence, suspension clamp position, suspension clamp confidence, strain clamp position, strain clamp confidence, vibration damper position, vibration damper confidence, grading ring position, grading ring confidence, fault heating point position, and fault heating point confidence in the infrared image of the transmission line to be detected; obtain the target detection result according to the insulator position, insulator confidence, suspension clamp position, suspension clamp confidence, strain clamp position, strain clamp confidence, vibration damper position, vibration damper confidence, grading ring position, grading ring confidence, fault heating point position, and fault heating point confidence in the infrared image of the transmission line to be detected.
[0152] Further, in some embodiments, the feature aggregation and detection module 300 is used to: perform adaptive operations on the enhanced global important information of the image based on multiple context enhancement modules to obtain multiple adaptive weights; calculate the weighted sum of the multiple adaptive weights to obtain the aggregated image information.
[0153] Further, in some embodiments, the improved YOLOv5s algorithm replaces the loss function of the original YOLOv5s algorithm with Wise-IoU v3. Before obtaining the infrared image of the transmission line to be detected, the acquisition module 100 is further configured to: obtain an infrared image set for training; perform labeling processing on the infrared image set to obtain a labeled data set, where the labels are in YOLO format and are divided into six labels: insulator, suspension clamp, strain clamp, vibration damper, grading ring, and fault heating point; based on a preset ratio, divide the labeled data set into a training set and a test set, and use the training set to train the improved YOLOv5s algorithm, and after the training is completed, based on Wise-IoU v3, use the test set to test the improved YOLOv5s algorithm to obtain a test result; if the test result meets the preset test conditions, use the improved YOLOv5s algorithm to perform multi-object detection on the infrared image of the transmission line to be detected, otherwise, adjust the preset ratio and retrain based on the adjusted training set until the new test result meets the preset test conditions.
[0154] It should be noted that the foregoing explanation of the embodiments of the multi-object detection method for infrared images of transmission lines also applies to the multi-object detection device for infrared images of transmission lines in this embodiment, and will not be repeated here.
[0155] According to the multi-object detection device for infrared images of transmission lines in the embodiments of the present invention, based on the Backbone part of the improved YOLOv5s algorithm, the attention weight feature information of the infrared image of the transmission line to be detected is calculated, and the global feature relationship information is obtained according to the initial high-dimensional information and the attention weight feature information, and the enhanced global important information of the image is obtained according to the processed global feature relationship information and the initial high-dimensional information of the infrared image of the transmission line to be detected; based on the Neck part of the improved YOLOv5s algorithm, feature aggregation is performed on the enhanced global important information of the image to obtain aggregated image information, and based on the Head part of the improved YOLOv5s algorithm, the target detection result is obtained according to the aggregated image information. Thereby, the problems of low detection accuracy and performance degradation in actual applications of traditional detection methods are solved, and the accuracy and robustness of target detection are improved.
[0156] Figure 16 The structural schematic diagram of the electronic device provided by the embodiments of the present invention. The electronic device may include:
[0157] A memory 1601, a processor 1602, and a computer program stored on the memory 1601 and executable on the processor 1602.
[0158] When the processor 1602 executes the program, it implements the multi-object detection method for infrared images of transmission lines provided in the foregoing embodiments.
[0159] Further, the electronic device further includes:
[0160] A communication interface 1603 for communication between the memory 1601 and the processor 1602.
[0161] A memory 1601 for storing a computer program that can run on the processor 1602.
[0162] The memory 1601 may include a high-speed RAM memory and may also include a non-volatile memory, such as at least one disk memory.
[0163] If the memory 1601, the processor 1602, and the communication interface 1603 are implemented independently, the communication interface 1603, the memory 1601, and the processor 1602 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 16 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0164] Optionally, in a specific implementation, if the memory 1601, the processor 1602, and the communication interface 1603 are integrated on a chip, the memory 1601, the processor 1602, and the communication interface 1603 can communicate with each other through an internal interface.
[0165] The processor 1602 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0166] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the multi-target detection method for infrared images of transmission lines as described above is implemented.
[0167] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0168] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0169] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A multi-target detection method for infrared images of transmission lines, characterized in that, The multi-object detection method for the infrared image of the transmission line uses an improved YOLOv5s algorithm. The improved YOLOv5s algorithm replaces the loss function of the original YOLOv5s algorithm with Wise-IoU v3. The improved YOLOv5s algorithm introduces a global context attention mechanism in the Backbone part of the original YOLOv5s algorithm and adds multiple context enhancement modules in the Neck part of the original YOLOv5s algorithm. The global context attention mechanism includes multiple C3GC modules, which are respectively set in the second, fourth, sixth, and eighth layers of the Backbone part. The C3GC module is obtained by embedding a GC module before one convolutional module in the C3 module and keeping the other convolutional module unchanged. The GC module is a global context module. The multiple context enhancement modules are respectively set before the second concat module to the fourth concat module in the Neck part. Among them, the method includes the following steps: Obtain the infrared image of the transmission line to be detected; Based on the Backbone part of the improved YOLOv5s algorithm, calculate the attention weight feature information of the infrared image of the transmission line to be detected, and obtain the global feature relationship information according to the initial high-dimensional information of the infrared image of the transmission line to be detected and the attention weight feature information, and process the global feature relationship information, and obtain the enhanced global important information of the image according to the processed global feature relationship information and the initial high-dimensional information of the infrared image of the transmission line to be detected; Based on the Neck part of the improved YOLOv5s algorithm, perform feature aggregation on the enhanced global important information of the image to obtain aggregated image information, and based on the Head part of the improved YOLOv5s algorithm, obtain the target detection result according to the aggregated image information.
2. The multi-object detection method for infrared images of transmission lines according to claim 1, characterized in that The target detection result includes at least one of the insulator position and insulator confidence, suspension clamp position and suspension clamp confidence, strain clamp position and strain clamp confidence, vibration damper position and vibration damper confidence, grading ring position and grading ring confidence, fault heating point position and fault heating point confidence.
3. The multi-object detection method for infrared images of transmission lines according to claim 2, characterized in that, The step of obtaining the target detection result according to the aggregated image information based on the Head part of the improved YOLOv5s algorithm includes: Use the Head part of the improved YOLOv5s algorithm to generate detection frames, and perform classification, localization, and confidence scoring based on the detection frames and the aggregated image information to obtain the insulator position, insulator confidence, suspension clamp position, suspension clamp confidence, strain clamp position, strain clamp confidence, vibration damper position, vibration damper confidence, grading ring position, grading ring confidence, fault heating point position, and fault heating point confidence in the infrared image of the transmission line to be detected; The target detection result is obtained according to the insulator position, insulator confidence, suspension clamp position, suspension clamp confidence, strain clamp position, strain clamp confidence, vibration damper position, vibration damper confidence, grading ring position, grading ring confidence, fault heating point position and fault heating point confidence in the infrared image of the transmission line to be detected.
4. The multi-object detection method for infrared images of transmission lines according to claim 1, characterized in that The Neck part of the improved YOLOv5s algorithm aggregates the global important information of the enhanced image to obtain aggregated image information, including: Based on the multiple context enhancement modules, adaptive operations are performed on the global important information of the enhanced image to obtain multiple adaptive weights; The weighted sum of the multiple adaptive weights is calculated to obtain the aggregated image information.
5. The multi-object detection method for infrared images of transmission lines according to claim 1, characterized in that Before obtaining the infrared image of the transmission line to be detected, it further includes: Obtaining an infrared image set for training; Performing a labeling process on the infrared image set to obtain a labeled data set, where the labels are in the YOLO format and are divided into six labels: insulator, suspension clamp, strain clamp, vibration damper, grading ring, and fault heating point; Based on a preset ratio, the labeled data set is divided into a training set and a test set, and the improved YOLOv5s algorithm is trained using the training set, and after training is completed, the improved YOLOv5s algorithm is tested using the test set based on the Wise-IoU v3 to obtain a test result; If the test result meets the preset test conditions, the improved YOLOv5s algorithm is used to perform multi-object detection on the infrared image of the transmission line to be detected, otherwise, the preset ratio is adjusted, and re-training is performed based on the adjusted training set until the new test result meets the preset test conditions.
6. A multi-object detection device for infrared images of transmission lines, characterized in that, The multi-object detection device for the infrared image of the transmission line uses the improved YOLOv5s algorithm. The improved YOLOv5s algorithm replaces the loss function of the original YOLOv5s algorithm with Wise-IoU v3. The improved YOLOv5s algorithm introduces a global context attention mechanism in the Backbone part of the original YOLOv5s algorithm and adds multiple context enhancement modules in the Neck part of the original YOLOv5s algorithm. The global context attention mechanism includes multiple C3GC modules, which are respectively set in the second, fourth, sixth, and eighth layers of the Backbone part. The C3GC module is obtained by embedding a GC module before one convolutional module in the C3 module and keeping the other convolutional module unchanged. The GC module is a global context module. The multiple context enhancement modules are respectively set before the second concat module to the fourth concat module in the Neck part. Among them, the device includes: An acquisition module for acquiring the infrared image of the transmission line to be detected; A processing module, configured to calculate attention weight feature information of the infrared image of the transmission line to be detected based on the Backbone part of the improved YOLOv5s algorithm, obtain global feature relationship information according to the initial high-dimensional information of the infrared image of the transmission line to be detected and the attention weight feature information, process the global feature relationship information, and obtain enhanced global important information of the image according to the processed global feature relationship information and the initial high-dimensional information of the infrared image of the transmission line to be detected; A feature aggregation and detection module, configured to perform feature aggregation on the enhanced global important information of the image based on the Neck part of the improved YOLOv5s algorithm to obtain aggregated image information, and obtain a target detection result based on the Head part of the improved YOLOv5s algorithm according to the aggregated image information.
7. The multi-object detection device for infrared images of transmission lines according to claim 6, characterized in that The target detection result includes at least one of the insulator position and insulator confidence, suspension clamp position and suspension clamp confidence, strain clamp position and strain clamp confidence, vibration damper position and vibration damper confidence, grading ring position and grading ring confidence, and fault heating point position and fault heating point confidence.
8. The multi-target detection device for infrared images of transmission lines according to claim 7, characterized in that, The feature aggregation and detection module is configured to: Generate detection boxes using the Head part of the improved YOLOv5s algorithm, and perform classification, localization, and confidence scoring based on the detection boxes and the aggregated image information to obtain the insulator position, insulator confidence, suspension clamp position, suspension clamp confidence, strain clamp position, strain clamp confidence, vibration damper position, vibration damper confidence, grading ring position, grading ring confidence, fault heating point position, and fault heating point confidence in the infrared image of the transmission line to be detected; Obtain the target detection result according to the insulator position, insulator confidence, suspension clamp position, suspension clamp confidence, strain clamp position, strain clamp confidence, vibration damper position, vibration damper confidence, grading ring position, grading ring confidence, fault heating point position, and fault heating point confidence in the infrared image of the transmission line to be detected.
9. An electronic device, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the multi-target detection method for the infrared image of the transmission line according to any one of claims 1-5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by the processor to be used to implement the multi-target detection method for the infrared image of the transmission line according to any one of claims 1-5.
Citation Information
Patent Citations
Aerial image rotating target detection method based on annular smooth label
CN116597324A