A multi-chamber lightning arrester recognition method based on improved YOLOX

CN115841608BActive Publication Date: 2026-09-29HAIBEI POWER SUPPLY COMPANY STATE GRID QINGHAI ELECTRIC POWER +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211364891.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-02
Publication Date
2026-09-29
Estimated Expiration
2042-11-02

AI Technical Summary

Technical Problem

手工选取特征费时费力,需要启发式专业知识,很大程度上依靠经验和运气

Benefits of technology

[0041]本发明提出将YOLOX目标检测算法与旋转框检测算法相结合,提高了电力系统中多腔室避雷器识别的检测精度,将YOLOX的主干网络替换为具有更大感受野的ConvNext,提高学习多腔室避雷器特征能力;在空间金字塔池化模块使用通道乱序操作增强特征融合,并增加旋转框检测思想以减少识别结果中的背景干扰。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115841608B_ABST
    Figure CN115841608B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of multi-chamber lightning arrester identification, and particularly relates to a multi-chamber lightning arrester identification method based on improved YOLOX, comprising the following steps: S1, using Labelme tool to label the multi-chamber lightning arrester in the image map; S2, improving the YOLOX algorithm, using a larger receptive field backbone network to replace DarkNet53, and using depth separable convolution to ensure the high speed of the model, using channel disorder operation in the spatial pyramid pooling module to increase the information interaction between the features, enhancing the feature fusion capability, and at the same time, adding the rotating frame detection idea to reduce the interference of the background in the identification result; S3, using the YOLOX target detection algorithm to realize the detection of the rotating target. The present application not only improves the detection precision of the multi-chamber lightning arrester identification in the power system, but also improves the learning ability of the multi-chamber lightning arrester characteristics, and further adds the rotating frame detection idea to reduce the background interference in the identification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-chamber surge arrester identification technology, and in particular to a multi-chamber surge arrester identification method based on an improved YOLOX. Background Technology

[0002] In power systems, the safety of transmission lines is paramount for the normal operation of the power grid, thus requiring regular inspections. For safety and stability, transmission lines are typically erected at high altitudes via towers. Multi-chamber surge arresters connect the power lines to the ground, protecting electrical equipment from high transient overvoltages and limiting their duration. During daily power system operation, surge arresters should be regularly inspected and maintained to ensure their proper functioning and prevent power line malfunctions. Traditional manual inspection methods are labor-intensive and inefficient. Currently, using drones to replace manual operations simplifies the work process and reduces the risks of aerial operations. A "drone-based, human-assisted" maintenance model has been largely established, continuously improving the level of transmission line maintenance.

[0003] Currently, the mainstream method for identifying surge arresters is vision-based online inspection using drones. The methods employed include spark gap detection, infrared thermal imaging, small ball discharge detection, leakage current detection, and laser Doppler vibration detection. Through computer data analysis and processing, the efficiency of surge arrester identification can be greatly improved, and in many advanced countries, it has gradually replaced traditional manual ground inspections.

[0004] In practical applications, the complexity of the circuit background, the varying installation locations, and the large number and diverse types of multi-chamber surge arresters present significant challenges for equipment identification. Furthermore, identification requires extensive manual feature extraction, as different research objects exhibit different features, resulting in feature diversity, such as SIFT, HOG, and LBP. Manual feature selection is time-consuming and labor-intensive, requiring heuristic expertise and relying heavily on experience and luck. Current models, such as SVM, Boosting, and LR, are shallow learning methods, with limited ability to represent complex functions given a limited number of samples and computational units. Their generalization ability for complex classification problems is limited, and some models, such as artificial neural networks (BP), are prone to getting trapped in local minima during training.

[0005] Therefore, this invention proposes a multi-chamber surge arrester identification method based on improved YOLOX, which can provide important data support for surge arrester detection. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of existing technologies by proposing an improved YOLOX-based method for identifying multi-chamber surge arresters.

[0007] A method for identifying multi-chamber surge arresters based on an improved YOLOX includes the following steps:

[0008] S1. Use the Labelme tool to label the multi-chamber surge arresters in the image;

[0009] S2. Improve the YOLOX algorithm by replacing DarkNet53 with a backbone network with a larger receptive field and using depthwise separable convolution to ensure the high speed of the model. In the spatial pyramid pooling module, channel disorder operation is used to increase the information interaction between features and enhance the feature fusion capability. At the same time, the idea of ​​rotating box detection is added to reduce the interference of background in the recognition results.

[0010] S3. Implement the detection of rotating targets using the YOLOX target detection algorithm, with the following specific improvements:

[0011] S3-1: Detecting horizontal rectangular boxes requires obtaining the x, y, w, and h coordinates of the box to represent it. However, for rotated rectangular boxes, the rotation angle θ is also needed. Therefore, a branch needs to be added to the network's detection head output to obtain the rotation angle θ. The circular smooth label (CSL) encoding method is used to classify the box angle, and its expression is as follows:

[0012]

[0013] In the formula, θ represents the rotation angle of the current ground truth bounding box, r represents the window radius (default value is 6), and g(x) represents the window function, the expression of which is as follows:

[0014]

[0015] Furthermore, the window function satisfies periodicity and symmetry, and its corresponding expressions are shown below:

[0016] g(x)=g(x+KT), K∈N, T=180 / ω,

[0017] 0≤g(θ+ε)=g(θ-ε)≤1,|ε|<r;

[0018] In the formula, the mean μ = 0, the variance δ = 4, and N is the set of natural numbers;

[0019] S3-2: In YOLOX, there are two places where the loss function needs to be calculated. The first loss calculation is used to construct the cost matrix for label assignment, and its loss function is as follows. The second loss calculation is only to filter positive and negative samples:

[0020] Cost = L cls +λL reg +L1,

[0021]

[0022] In the formula, the classification loss L cls The binary cross-entropy loss function is used, and the regression loss L... reg The IOU loss function is used, and λ is used to control the ratio of the two loss weights. λ defaults to 3. L1 in the formula is equivalent to adding more prior information and increasing the matching degree of high-quality priors.

[0023] The second loss calculation is used to optimize the model, and its loss function is as follows:

[0024] Loss = L cls +λL reg +L conf ;

[0025] The formula includes classification loss L cls Confidence loss L conf and regression loss L reg The classification loss and confidence loss both use the binary cross-entropy loss function, with λ controlling the regression loss weight ratio (default is 5). The true labels for confidence and classification are swapped. Since the improved algorithm has an additional branch for rotation angle classification in its output, the loss function in YOLOX needs to be improved.

[0026] The improved cost function for label allocation is as follows:

[0027] Cost = L cls +λL reg +L1+L θ-cls ;

[0028] θ-cls is used to calculate the angle classification loss, and the angle loss is calculated using the sigmoid combined with the binary cross-entropy.

[0029] The loss function during model training is improved as follows:

[0030] Loss = L cls +λL reg +L conf +L θ-cls ;

[0031] The optimized loss function adds an angle classification loss to the original loss, and the angle loss is calculated using the same sigmoid function combined with binary cross-entropy.

[0032] As a preferred embodiment of the present invention, the labeling method for experimental data in S1 is divided into two modes: rectangular box labeling and rotated rectangular box labeling.

[0033] As a preferred embodiment of the present invention, step S2 includes the following sub-steps:

[0034] S2-1: To improve the model's feature learning ability, the original backbone network in YOLOX is improved using the basic units and downsampling units in the ConvNext module, which have a larger receptive field, as follows:

[0035] The input features are first extracted by a 7×7 depthwise separable convolution. The larger the convolution kernel, the wider the feature region learned. Then, two 1×1 ordinary convolution operations are used to increase and decrease the dimensionality of the features to obtain more abstract semantic information. Finally, the features are added and fused with the features of the residual mapping branch.

[0036] Downsampling can reduce the resolution of input features and reduce the computational complexity of the model. A 3×3 depthwise separable convolution is added to the residual mapping, and the stride is set to 2 with the 7×7 depthwise separable convolution to reduce the feature resolution by half and increase the number of channels by half.

[0037] S2-2: The backbone network structure after improving the feature extraction network using the ConvNet module is 640×640×3. In the first-level feature, one basic unit and one downsampling unit are used to learn shallow features, and the feature resolution is reduced by half while the channel dimension is doubled. Similarly, in the second-level feature, two basic units and one downsampling unit are used to learn features. In the third-level feature, five basic units and one downsampling unit are used to learn features. In the fourth-level feature, two basic units and one downsampling unit are used to learn features. In the fifth-level feature, one basic unit and one downsampling unit are used to learn features. Finally, the improved SPP module is used to obtain multi-scale features.

[0038] S2-3: First, four sets of max pooling operations of different sizes are applied to the input features to obtain four sets of compressed feature vectors. In order to reduce the number of parameters in the module and to ensure that the feature dimension of each set of features is consistent with the input when concatenating, a 1×1 convolution operation is used to compress its channel dimension by a factor of 4. Then, upsampling is used to restore the feature resolution, and the four sets of features are merged in the channel dimension. The merged result is added to the original input features and fused. Finally, in order to increase the information transfer between features, a channel disorder operation is used to reorder the final result in the channel dimension.

[0039] As a preferred embodiment of the present invention, the rotation angle θ in S3 is in the range of 0°≤θ<180°.

[0040] Compared with the prior art, the beneficial effects of the present invention are:

[0041] This invention proposes to combine the YOLOX target detection algorithm with the rotating frame detection algorithm to improve the detection accuracy of multi-chamber surge arresters in power systems. The backbone network of YOLOX is replaced with ConvNext, which has a larger receptive field, to improve the ability to learn the features of multi-chamber surge arresters. In the spatial pyramid pooling module, channel disorder operation is used to enhance feature fusion, and the rotating frame detection idea is added to reduce background interference in the recognition results. Attached Figure Description

[0042] Figure 1 This is a schematic diagram of the method flow for identifying multi-chamber surge arresters based on an improved YOLOX proposed in this invention;

[0043] Figure 2 This is a schematic diagram of two labeling methods for a multi-chamber surge arrester identification method based on the improved YOLOX proposed in this invention;

[0044] Figure 3 This is a schematic diagram of the original backbone network of YOLOX in the ConvNext module of the improved YOLOX-based multi-chamber surge arrester identification method proposed in this invention, which includes the basic unit and downsampling unit.

[0045] Figure 4 This is a schematic diagram of the improved SPP structure for a multi-chamber surge arrester identification method based on improved YOLOX proposed in this invention;

[0046] Figure 5 This is a schematic diagram of the CSL-coded tag score for a multi-chamber surge arrester identification method based on the improved YOLOX proposed in this invention. Detailed Implementation

[0047] The present invention will be further explained below with reference to specific embodiments.

[0048] Reference Figure 1-5 A method for identifying multi-chamber surge arresters based on an improved YOLOX includes the following steps:

[0049] S1. Use the Labelme tool to label the multi-chamber surge arresters in the image;

[0050] The aerial imagery includes 2000 training images and 500 test images. For tilted targets, if horizontal bounding boxes are used for detection, these boxes will contain a large amount of background information. Therefore, this invention uses two annotation modes for the experimental data: bounding box annotation and rotated bounding box annotation. The results are as follows: Figure 2 As shown, where Figure 2 -a indicates the result of the rectangular bounding box annotation, which includes more of the background. Figure 2 -b indicates the result of rotating the bounding box annotation;

[0051] Given the difficulty in acquiring image data and the tedious and time-consuming annotation process, and considering that neural network models require a large amount of data for fitting during the learning of target features, data augmentation processing is performed on the training set. Multi-chamber surge arresters have orientation feature invariance in spatial location, scale variability in shape, and are subject to occlusion at shooting angles. Here, data augmentation methods such as mirror enhancement, multi-scale scaling, and random erasure are used to expand the training data to 3000 images.

[0052] S2. Improve the YOLOX algorithm by replacing DarkNet53 with a backbone network having a larger receptive field, and using depthwise separable convolutions to ensure the model's high speed. In the spatial pyramid pooling module, channel disorder operations are used to increase the information interaction between features and enhance feature fusion capabilities. Simultaneously, the idea of ​​rotated bounding box detection is added to reduce background interference in the recognition results. This step includes the following sub-steps:

[0053] S2-1: To improve the model's feature learning ability, the original backbone network in YOLOX is improved using the Base Unit (BU) and Down Sample Unit (DSU) from the ConvNext module, which have a larger receptive field. The structure is as follows: Figure 3 As shown, the details are as follows:

[0054] Figure 3 -a is the BU module. The input features are first extracted by a 7×7 depthwise separable convolution. The larger the convolution kernel, the wider the feature region learned. Then, two 1×1 ordinary convolution operations are used to increase and decrease the dimensionality of the features to obtain more abstract semantic information. Finally, the features are added and fused with the features of the residual mapping branch.

[0055] Figure 3 -b shows the DSU module. Downsampling can reduce the resolution of input features and reduce the computational complexity of the model. A 3×3 depthwise separable convolution is added to the residual mapping, and the stride is set to 2 with the 7×7 depthwise separable convolution to reduce the feature resolution by half and increase the number of channels by half.

[0056] S2-2: The backbone network structure after improving the feature extraction network using the ConvNet module is shown in the table below. The network input size is 640×640×3. In the first-level feature, one BU module and one DSU module are used to learn shallow features, and the feature resolution is reduced by half while the channel dimension is doubled. Similarly, in the second-level feature, two BU modules and one DSU module are used to learn features. In the third-level feature, five BU modules and one DSU module are used to learn features. In the fourth-level feature, two BU modules and one DSU module are used to learn features. In the fifth-level feature, one BU module and one DSU module are used to learn features. Finally, the improved SPP module is used to obtain multi-scale features.

[0057]

[0058]

[0059] S2-3: The SPP module uses multiple sets of max pooling operations to compress features to a fixed size, solving the problem of fixed image input size caused by the model structure, while introducing multi-scale features to the model. This invention makes adaptive modifications to SPP, and its structure is as follows: Figure 4 As shown, firstly, four sets of max pooling operations of different sizes are applied to the input features to obtain four sets of compressed feature vectors. In order to reduce the number of parameters of the module and to ensure that the feature dimension of each set of features is consistent with the input when concatenating, a 1×1 convolution operation is used to compress its channel dimension by a factor of 4. Then, upsampling is used to restore the feature resolution, and the four sets of features are merged in the channel dimension. The merged result is added to the original input features and fused. Finally, in order to increase the information transfer between features, a channel disorder operation is used to reorder the final result in the channel dimension.

[0060] S3. Implement the detection of rotating targets using the YOLOX target detection algorithm, with the following specific improvements:

[0061] S3-1: Detecting a horizontal rectangle requires obtaining its x, y, w, and h coordinates to represent it. However, for rotated rectangles, the rotation angle θ is also needed. Therefore, a branch needs to be added to the network's detection head output to obtain the rotation angle θ. This invention defines a rotated rectangle using a long-side representation, in the format of x, y, w, h, and θ (0° ≤ θ < 180°). θ is not obtained through regression because the periodicity of the angle would interfere with the regression results. Instead, a classification task is used to obtain the angle θ. However, due to the periodicity of the angle, simple one-hot encoding cannot be used. Instead, a Circular Smooth Label (CSL) encoding method is used to classify the box angle. Its expression is as follows:

[0062]

[0063] In the formula, θ represents the rotation angle of the current ground truth bounding box, r represents the window radius (default value is 6), and g(x) represents the window function, the expression of which is as follows:

[0064]

[0065] Furthermore, the window function satisfies periodicity and symmetry, and its corresponding expressions are shown below:

[0066] g(x)=g(x+KT), K∈N, T=180 / ω,

[0067] 0≤g(θ+ε)=g(θ-ε)≤1,|ε|<r;

[0068] In the formula, the mean μ = 0, the variance δ = 4, N is the set of natural numbers, ε also represents the angle, 0 ≤ g(θ + ε) = g(θ - ε) ≤ 1, |ε| < r. This formula indicates that the window function satisfies symmetry.

[0069] When the actual θ value is 0° or 90°, the corresponding CSL encoded label scores are as follows: Figure 5 As shown in the left and right images, the horizontal axis represents the angle value and the vertical axis represents the coded label score.

[0070] S3-2: In YOLOX, there are two places where the loss function needs to be calculated. The first loss calculation is used to construct the cost matrix for label assignment, and its loss function is as follows. The second loss calculation is only to filter positive and negative samples:

[0071] Cost = L cls +λL reg +L1,

[0072]

[0073] In the formula, the classification loss L cls The binary cross-entropy loss function is used, and the regression loss L... reg The IOU loss function is used, and λ is used to control the weight ratio of the two losses. λ defaults to 3. L1 in the formula is equivalent to adding more prior information and increasing the matching degree of high-quality priors (the center point of the grid is within the range of the true box or within 2.5 grids from the center of the true box).

[0074] The second loss calculation is used to optimize the model, and its loss function is as follows:

[0075] Loss = L cls +λL reg +L conf ;

[0076] The formula includes classification loss L cls Confidence loss L conf and regression loss L reg The classification loss and confidence loss both use the binary cross-entropy loss function, with λ controlling the regression loss weight ratio (default is 5). The true labels for confidence and classification are swapped. Since the improved algorithm has an additional branch for rotation angle classification in its output, the loss function in YOLOX needs to be improved.

[0077] The improved cost function for label allocation is as follows:

[0078] Cost = L cls +λL reg +L1+L θ-cls ;

[0079] θ-cls is used to calculate the angle classification loss, and the angle loss is calculated using the sigmoid combined with the binary cross-entropy form.

[0080] The loss function during model training is improved as follows:

[0081] Loss = L cls +λL reg +L conf +L θ-cls ;

[0082] The optimized loss function adds an angle classification loss to the original loss, and the angle loss is calculated using the same sigmoid function combined with binary cross-entropy.

[0083] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for identifying multi-chamber surge arresters based on an improved YOLOX, characterized in that, Includes the following steps: S1. Use the Labelme tool to label the multi-chamber surge arresters in the aerial images; the aerial images include 2000 training images and 500 test images. S2. Improve the YOLOX algorithm by replacing DarkNet53 with a backbone network having a larger receptive field and using depthwise separable convolutions; use channel disordered operations in the spatial pyramid pooling module, and add the idea of ​​rotated bounding box detection; S2 includes the following sub-steps: S2-1: Improve the original backbone network in YOLOX using the basic units and downsampling units in the ConvNext module, which has a larger receptive field, as follows: In the basic unit, the input features are first extracted by a 7×7 depthwise separable convolution, then two 1×1 ordinary convolutions are used to perform feature dimensionality upscaling and downscaling, and finally the features are added and fused with the features of the residual mapping branch. In the downsampling unit, the input features are first extracted by a 7×7 depthwise separable convolution with a stride of 2; then, two 1×1 ordinary convolutions are used to perform feature upscaling and downscaling operations, and finally, the features are added and fused with the features of the residual mapping branch, where a 3×3 depthwise separable convolution is added to the residual mapping. S2-2: The backbone network structure of the feature extraction network is improved using the ConvNext module. The network input size is 640×640×3. In the first-level feature, one basic unit and one downsampling unit are used to learn shallow features, and the feature resolution is reduced by half while the channel dimension is doubled. In the second-level feature, two basic units and one downsampling unit are used to learn features. In the third-level feature, five basic units and one downsampling unit are used to learn features. In the fourth-level feature, two basic units and one downsampling unit are used to learn features. In the fifth-level feature, one basic unit and one downsampling unit are used to learn features. Finally, the improved SPP module is used to obtain multi-scale features. S2-3: First, the input features are subjected to four sets of max pooling operations of different sizes to obtain four sets of compressed feature vectors. Then, a 1×1 convolution operation is used to compress the channel dimension by a factor of 4. Then, upsampling is used to restore the feature resolution, and the four sets of features are merged in the channel dimension. The merged result is added to the original input features and fused. Finally, a channel disorder operation is used to reorder the final result in the channel dimension. S3. Implement the detection of rotating targets using the YOLOX target detection algorithm, with the following specific improvements: S3-1: The detection of horizontal rectangular boxes is represented by four parameters: X, Y, W, and H. The detection of rotated rectangular boxes is represented by five parameters: X, Y, W, H, and rotation angle θ. A branch is added to the output of the network's detection head to obtain the rotation angle θ. The circular smooth label (CSL) encoding method is used to classify the angle of the boxes. The expression is as follows: ; In the formula, θ represents the rotation angle of the current ground truth bounding box, r represents the window radius, which takes the value of 6, and g(x) represents the window function, the expression of which is as follows: ; Furthermore, the window function satisfies periodicity and symmetry, and its corresponding expressions are as follows: ; ; In the formula, the mean ,variance , It is the set of natural numbers; S3-2: YOLOX includes two loss function calculations. The first loss calculation is used to construct the cost matrix for label assignment, and its loss function is as follows. The second loss calculation is used to filter positive and negative samples: ; In the formula, the classification loss L cls The binary cross-entropy loss function is used, and the regression loss L... reg The IOU loss function is used, and λ1 is used to control the ratio of the two loss weights, with a value of 3. When the center point of the grid is within the range of the true bounding box or within 2.5 grids from the center of the true bounding box, L1 is -1000, otherwise it is 0. The second loss calculation is used to optimize the model, and its loss function is as follows: ; The formula includes classification loss L cls Confidence loss L conf and regression loss L reg The classification loss and confidence loss both use the binary cross-entropy loss function, with λ2 controlling the regression loss weight ratio and set to 5. The true labels for confidence and classification are swapped. The improved algorithm adds a rotation angle classification branch to its output branch. The loss function in YOLOX is also improved. The improved cost function for label allocation is as follows: ; Where L θ-cls To calculate the angle classification loss, the sigmoid function combined with binary cross-entropy is used to calculate the angle loss. The loss function during model training is improved as follows: ; The optimized loss function adds an angle classification loss to the original loss, and the angle loss is calculated using the same sigmoid function combined with binary cross-entropy.

2. The method for identifying multi-chamber surge arresters based on the improved YOLOX according to claim 1, characterized in that, The labeling method for experimental data in S1 is divided into two modes: rectangular box labeling and rotated rectangular box labeling.

3. The method for identifying multi-chamber surge arresters based on the improved YOLOX according to claim 1, characterized in that, The range of the rotation angle θ in S3 is 0°≤θ<180°.

Citation Information

Patent Citations

  • Unmanned aerial vehicle aerial photography target detection and identification method and system

    CN114842365A

  • SAR (Synthetic Aperture Radar) image target detection method based on full-space coding attention module

    CN115147731A