Insulator ultraviolet corona discharge target detection method based on YOLO-SM

By combining the YOLO-SM network with SCCDG and MSDAM modules, feature extraction and loss functions are optimized, solving the problems of small target accuracy and complex background recognition in corona discharge detection, and achieving efficient and accurate corona discharge target detection.

CN120912997AActive Publication Date: 2025-11-07NANCHANG POWER SUPPLY BRANCH OF STATE GRID JIANGXI ELECTRIC POWER CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511450846.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2025-11-07
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

Existing corona discharge detection methods suffer from insufficient accuracy in detecting small targets, limited recognition capabilities in complex backgrounds, and poor adaptability to irregularly shaped targets, resulting in low detection accuracy and high false detection rates, making it difficult to meet the needs of practical engineering applications.

Method used

A detection method based on YOLO-SM is adopted, which combines the lightweight feature extraction module SCCDG and the multi-scale dual attention mechanism module MSDAM. The detection network is optimized by a composite bounding box regression loss function to achieve multi-scale feature extraction and attention-weighted fusion, thereby enhancing the recognition ability and localization accuracy of small targets.

Benefits of technology

It significantly improves the accuracy and recall rate of small target detection, reduces the false alarm rate and false alarm rate, and improves the detection speed and computational efficiency, making it suitable for real-time inspection deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912997A_ABST
    Figure CN120912997A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and power equipment detection, in particular to an insulator ultraviolet corona discharge target detection method based on YOLO-SM. The method comprises the steps of constructing a feature extraction network integrated with an SCCDG lightweight feature extraction module, a feature fusion network of an MSDAM multi-scale double attention mechanism module and a prediction network, and adopting composite bounding box regression loss training. The SCCDG lightweight feature extraction module is used for averagely separating input channels and sending sub-features into CDG multi-channel convolution so as to extract multi-scale features; generating attention weights through an MSDAM multi-scale double attention mechanism module, and carrying out weighted fusion on the attention weights; the prediction network is provided with a multi-scale detection head, and a detection result is output through confidence threshold and non-maximum suppression. The method gives consideration to the detection precision and the calculation efficiency, is suitable for real-time monitoring scenes such as unmanned aerial vehicle inspection, and facilitates the improvement of the detection rate and the positioning precision of the tiny corona discharge target.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the field of artificial intelligence and power equipment detection, and particularly relates to an insulator ultraviolet corona discharge target detection method based on YOLO-SM. BACKGROUND

[0002] An insulator is a key component of a power transmission line, and an ultraviolet corona discharge phenomenon of the insulator is an important precursor of a line fault. Timely and accurate detection of the corona discharge is of great significance to guarantee safe and stable operation of a power system. With the development of artificial intelligence technology, a target detection method based on deep learning provides a new technical approach for automatic identification of the corona discharge.

[0003] The existing corona discharge detection methods have many technical defects: a traditional YOLO series algorithm has insufficient feature extraction capability when processing small target corona discharge, resulting in low detection precision; although an existing lightweight network has high calculation efficiency, the recognition capability of the lightweight network for weak corona signals in a complex background is limited, and the lightweight network is prone to missed detection and false detection; an existing attention mechanism mostly adopts a single scale design, and cannot effectively fuse multi-scale feature information, and has poor adaptability to irregularly shaped corona discharge targets; and a traditional bounding box regression loss function is mainly based on IoU calculation, and cannot accurately fit the irregular boundary of a corona discharge target, and has low positioning precision. These technical limitations seriously restrict the application effect of a corona discharge detection system in actual engineering. SUMMARY

[0004] The application provides an insulator ultraviolet corona discharge target detection method based on YOLO-SM, which aims to improve small target detection precision, suppress background noise, optimize positioning loss, reduce calculation overhead, and support real-time inspection deployment.

[0005] The insulator ultraviolet corona discharge target detection method based on YOLO-SM comprises the following steps: S100: obtaining a trained YOLO-SM detection network, the network comprising a feature extraction network, a feature fusion network and a prediction network, wherein: The feature extraction network is integrated with an SCCDG lightweight feature extraction module, the SCCDG lightweight feature extraction module separates multiple sub-features through an average separation channel operation, and performs three-way processing of ordinary convolution, dilated convolution and grouped convolution on the input features in parallel through a CDG module; The feature fusion network is integrated with an MSDAM multi-scale dual attention mechanism module, the module comprising two parallel branches, and the two parallel branches adopt different size pooling kernels and are respectively processed by Mish and Swish activation functions to generate complementary attention weights; The YOLO-SM detection network is trained using a composite bounding box regression loss function; S200: input the insulator ultraviolet corona discharge image into the YOLO-SM detection network, pass the input features through the SCDDG lightweight feature extraction module integrated by the feature extraction network for channel splitting and multi-branch convolution processing, and extract multi-scale features; S300: input the multi-scale features into the feature fusion network, pass the multi-scale features through the double-branch structure of the MSDAM multi-scale double attention mechanism module for different scale pooling operations respectively, generate attention weights, and perform weighted fusion on the multi-scale features to obtain fused features; S400: input the fused features into the prediction network for target detection, and output the bounding box coordinates, class labels and confidence of the corona discharge target.

[0006] As a preferred technical solution of the present application, the composite bounding box regression loss function includes the weighted sum of the following four losses: The intersection over union loss is calculated by the loss value of the overlap of the prediction box and the target box; The shape deviation loss is calculated by the difference in aspect ratio between the prediction box and the target box; The corner distance loss is calculated by the distance between the corner coordinates of the prediction box and the target box; The density weighted loss sets a weighting factor based on the target distribution density.

[0007] As a preferred technical solution of the present application, the feature extraction network comprises a convolution module and a SCDDG lightweight feature extraction module, the convolution module is composed of two-dimensional convolution, batch normalization and activation function, and the outputs of multiple SCDDG lightweight feature extraction modules are connected to the feature fusion network to provide multi-level features.

[0008] As a preferred technical solution of the present application, the SCDDG lightweight feature extraction module comprises: The input features are separated into multiple sub-features by average separation channel operation, the sub-features are stacked after convolution operation, the stacked features are input into the CDG module for processing, the CDG module outputs are connected to the stacked features to form a residual connection, and finally the processed features are output after convolution operation; The CDG module performs three-way convolution processing on the input features in parallel: the first way performs ordinary convolution operation, the second way performs dilated convolution operation, and the third way performs grouped convolution operation; the three convolution results are stacked in the channel, and after convolution, normalization and activation function processing, convolution and pooling operations are performed, and a residual connection is formed with the intermediate processing result.

[0009] As a preferred technical solution of the present application, the feature fusion network comprises an upsampling module, a channel superposition module, a convolution module, an SCCDG lightweight feature extraction module and an MSDAM multi-scale dual attention mechanism module, the feature scale is adjusted through upsampling and downsampling operations, feature splicing is realized through channel superposition, and the multiple MSDAM multi-scale dual attention mechanism modules are respectively connected with detection heads of different scales.

[0010] As a preferred technical solution of the present application, the MSDAM multi-scale dual attention mechanism module comprises: after the input features are preprocessed, the input features are respectively input into two parallel branches, the first branch is processed through a first activation function after average pooling and maximum pooling operations of a first scale are adopted, the second branch is processed through a second activation function after average pooling and maximum pooling operations of a second scale are adopted, the two branches respectively form residual connections with the input features to generate attention weights, and finally weighted fusion is performed.

[0011] As a preferred technical solution of the present application, the first branch and the second branch adopt different sizes of pooling kernels for pooling operations, the first branch is processed through a Mish activation function after being pooled, the second branch is processed through a Swish activation function after being pooled, and different attention weight features are generated by the two branches.

[0012] As a preferred technical solution of the present application, the prediction network comprises three detection heads, each detection head comprises a bounding box regression branch, a class classification branch and a confidence prediction branch, and the three detection heads respectively process large-scale, medium-scale and small-scale target detection tasks.

[0013] As a preferred technical solution of the present application, the output results of the three detection heads are merged and then post-processed, including setting a confidence threshold to filter low-confidence detection boxes, removing redundant detection boxes through an IoU threshold by using a non-maximum suppression algorithm, and outputting final corona discharge target detection results.

[0014] The present application has the following advantages: 1. The present application realizes the best balance between feature extraction capability and computational efficiency through the channel separation-multibranch convolution design of the SCCDG lightweight feature extraction module and the dual-branch architecture of the MSDAM multi-scale dual attention mechanism module. The SCCDG lightweight feature extraction module processes three parallel paths through ordinary convolution, dilated convolution and grouped convolution, which can capture local details, expand the receptive field and reduce computational complexity; the MSDAM multi-scale dual attention mechanism module processes two branches through different scale pooling kernels and different activation functions, realizes complementary fusion of local detail features and global context information, and significantly enhances the recognition ability of weak corona signals in complex backgrounds.

[0015] 2. This invention integrates four dimensions—cross-union ratio loss, shape deviation loss, corner distance loss, and density-weighted loss—to form a composite bounding box regression loss function, systematically solving the detection challenge of irregularly shaped and blurred-boundary corona discharge targets. This loss function optimizes aspect ratio fitting through shape deviation loss, achieves refined boundary localization through corner distance loss, and ensures balanced learning in dense regions through density-weighted loss. Together with the aforementioned modules, it forms an end-to-end collaborative optimization system, effectively overcoming the technical bottlenecks of low detection accuracy and inaccurate localization of irregular corona discharge targets using traditional methods. Attached Figure Description

[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating the insulator ultraviolet corona discharge target detection method based on YOLO-SM of the present invention; Figure 2 This is the YOLO-SM network model structure of the present invention; Figure 3 This is a schematic diagram of the SCCDG lightweight feature extraction module in the YOLO-SM-based insulator ultraviolet corona discharge target detection method of the present invention; Figure 4 This is a schematic diagram of the CDG structure in the YOLO-SM-based insulator ultraviolet corona discharge target detection method of the present invention; Figure 5 This is a schematic diagram of the MSDAM multi-scale dual attention mechanism module in the YOLO-SM-based insulator ultraviolet corona discharge target detection method of the present invention. Detailed Implementation

[0017] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0018] Example 1: As Figure 1 As shown, the present invention provides a method for detecting ultraviolet corona discharge targets in insulators based on YOLO-SM, comprising: S100: Obtain the trained YOLO-SM detection network, which includes a feature extraction network, a feature fusion network, and a prediction network, wherein: The feature extraction network integrates a lightweight SCCDG feature extraction module. The lightweight SCCDG feature extraction module separates multiple sub-features through average channel separation operation, and performs three parallel processing steps on the input features through the CDG module: ordinary convolution, dilated convolution, and group convolution. The feature fusion network integrates an MSDAM multi-scale dual attention mechanism module, which includes two parallel branches that respectively use different size pooling kernels and are respectively processed by Mish and Swish activation functions to generate complementary attention weights; The YOLO-SM detection network is trained using a composite bounding box regression loss function; Further, the composite bounding box regression loss function includes a weighted sum of the following four losses: The intersection over union loss is calculated by the loss value of the overlap of the prediction box and the target box; The shape deviation loss is calculated by the difference in the aspect ratio of the prediction box and the target box; The corner point distance loss is calculated by the distance between the corner coordinates of the prediction box and the target box; The density weighted loss sets a weighting factor based on the target distribution density.

[0019] Specifically, a trained YOLO-SM detection network is obtained, as shown in Figure 2 The network uses an improved YOLO architecture and is specially optimized for insulator ultraviolet corona discharge small target detection tasks. The overall structure of the YOLO-SM detection network includes three main parts: a feature extraction network, a feature fusion network, and a prediction network. The feature extraction network is responsible for extracting multi-level deep features from the input 640x640x3 ultraviolet corona discharge image, including 5 ordinary convolution modules and 12 SCCDG lightweight feature extraction modules. Through layer-by-layer down-sampling, the input image is gradually reduced from 640x640 to 20x20, and the number of channels is increased from 3 to 1024. The feature fusion network realizes multi-scale feature fusion through up-sampling, down-sampling, and channel concatenation operations, and integrates an MSDAM multi-scale dual attention mechanism module to enhance target region feature expression. The prediction network includes three detection heads of different scales, which are used to detect large, medium, and small scale corona discharge targets.

[0020] The training of the YOLO-SM detection network uses a specially designed composite bounding box regression loss function, which is optimized for the characteristics of irregular shape and fuzzy boundary of corona discharge targets. The composite bounding box regression loss function is composed of the weighted sum of four loss terms: The first term is the intersection over union loss. IoU is the core indicator for evaluating the overlap of bounding boxes, but for irregularly shaped targets, using only the intersection over union loss cannot fully optimize the model. To improve accuracy, 1-IoU is introduced as a loss term to encourage the model to reduce positioning errors by increasing the overlap of bounding boxes. The loss function will prioritize optimizing the overlap between the prediction box and the target box during training to ensure that the prediction box encloses the target as much as possible. The 1-IoU formula is as follows: ; wherein: is the IoU loss value, and the value range is [0, 1]; 1-IoU represents the loss value of the overlap degree of the prediction box and the target box, and the greater the loss value, the smaller the overlap degree of the prediction box and the target box; is the area of the prediction box; is the area of the target box; represents the area of the overlapping region of the prediction box and the target box; represents the total area of the merging region of the prediction box and the target box.

[0021] When , it means that the prediction box and the target box are completely coincident; when 1, it means that the two boxes are completely not overlapped. In the back propagation process, the gradient of this loss term will drive the prediction box to move and adjust the size towards the target box to maximize the overlapping area.

[0022] In deep learning training, the goal of the optimizer is to minimize the loss function. If IoU is directly used as the loss, the greater the IoU (the higher the overlap degree), the greater the loss, and the optimizer will incorrectly reduce the overlap degree. After using 1-IoU, when the IoU increases, 1-IoU decreases, which conforms to the optimization logic that "loss reduction means performance improvement".

[0023] The second term is the shape deviation loss. The corona discharge target often presents an irregular shape, and its aspect ratio is an important morphological feature. The shape deviation loss constrains the difference in the aspect ratio of the prediction box and the target box, so that the model can more accurately capture the morphological features of the target. The calculation formula is: ; wherein: is the shape deviation loss value; is the width of the prediction box; is the height of the prediction box; is the width of the target box; is the height of the target box; is the number of samples, i.e. the number of target boxes.

[0024] The shape deviation loss and the IoU loss are complementary: the IoU loss focuses on the overall overlap degree, which may allow a large shape difference but still have a high IoU; the shape deviation loss strictly constrains the aspect ratio to ensure that the prediction box not only has the correct position, but also has the same shape as the target, which is particularly important for accurate detection of irregular corona discharge targets.

[0025] The third term is the corner distance loss. Traditional bounding box regression loss is difficult to accurately capture the coordinate positions of the four corners of the target box, especially when dealing with complex-shaped targets. The corner distance loss calculates the distance difference between the predicted box and the target box corner, achieving more precise positioning adjustment, improving the accuracy of bounding box prediction, and being suitable for asymmetric or angle complex scenes, effectively enhancing the model's perception of target boundary details and improving overall detection performance. The formula is as follows: ; Wherein: is the corner distance loss value; is the coordinate of the top-left corner of the predicted box; is the coordinate of the top-left corner of the target box; is the coordinate of the bottom-right corner of the predicted box; is the coordinate of the bottom-right corner of the target box; superscript represents top-left (top-left corner); superscript represents bottom-right (bottom-right corner); is the number of samples, i.e. the number of target boxes.

[0026] This loss term directly optimizes the corner position and is particularly effective for asymmetric or tilted targets. Compared with methods that only optimize the center point and width-height (such as IoU loss), the corner distance loss can more accurately constrain the spatial position of the bounding box. For example, two predicted boxes with the same IoU may have different corner deviations: one may be shifted as a whole, and the other may be distorted in shape. The corner distance loss can distinguish between the two cases and ensure that the four corners are accurately aligned, which is crucial for detecting boundary ambiguous and complex-shaped corona discharge targets.

[0027] The fourth term is the density weighted loss. In dense target scenes, the model tends to focus too much on isolated targets and ignore the details of dense areas. The density weighted loss assigns each target a weight factor based on the density of its surrounding targets, achieving balanced optimization in the training process. The calculation formula is as follows: ; Wherein: is the density weighted loss value; : is the weighted factor related to the target box, which depends on the density of the target box. Higher density areas get higher weights. : is the center coordinate of the predicted box; : is the center coordinate of the target box; : is the width of the target box; : is the height of the target box; : is the number of samples, i.e. the number of target boxes.

[0028] The density weighting factor ensures that the target in the dense area obtains a higher learning weight. For example, the weight of an isolated target is 1.0, and the weight of a target in a dense area surrounded by 5 adjacent targets is 1.5, so as to balance the learning effect of different distribution density areas. This design increases the loss contribution of the target in the dense area by 50%, forcing the model to pay more attention during the training process, thereby balancing the learning effect of different distribution density areas. For the scene of densely arranged multiple strings of insulators in the power transmission line, this loss term can effectively prevent the model from ignoring the corona discharge target in the dense area, significantly reducing the miss rate.

[0029] The four loss terms are combined by weighted summation to form the final composite bounding box regression loss function: ; wherein, is the total loss of the composite bounding box regression; is the weight factor, preferably 0.4, 0.2, 0.2, 0.2, used to balance the contribution of each loss term.

[0030] The synergistic mechanism of each loss term: the IoU loss provides an overall optimization direction, ensuring that the predicted box approaches the target box; the shape deviation loss further constrains the aspect ratio based on the IoU loss, preventing shape distortion; the corner distance loss achieves fine positioning, ensuring accurate alignment of the four corners; and the density weighting loss balances learning in different density areas, preventing training bias. This composite loss function simultaneously optimizes position, shape, corner, and density during the training process, forming a synergistic optimization mechanism that significantly improves the detection accuracy of irregular corona discharge targets.

[0031] S200: input the insulator ultraviolet corona discharge image into the YOLO-SM detection network, and pass the input features through the SCDDG lightweight feature extraction module integrated by the feature extraction network for channel splitting and multi-branch convolution processing to extract multi-scale features; Further, the feature extraction network comprises a convolution module and a SCDDG lightweight feature extraction module, the convolution module is composed of two-dimensional convolution, batch normalization and activation function, and the outputs of a plurality of SCDDG lightweight feature extraction modules are connected to a feature fusion network to provide multi-level features.

[0032] Further, the SCDDG lightweight feature extraction module comprises: The input features are separated into a plurality of sub-features by an average separation channel operation, the sub-features are stacked after convolution operation, the stacked features are input into a CDG module for processing, the output of the CDG module is connected to the stacked features in a residual connection, and finally the processed features are output after convolution operation; The CDG module performs three-way convolution processing on the input features in parallel: the first way performs ordinary convolution operation, the second way performs hollow convolution operation, and the third way performs grouping convolution operation; the three-way convolution results are stacked in channels, processed through convolution, normalization and activation function, and then through convolution and pooling operation, and connected with the intermediate processing results to form a residual connection.

[0033] Specifically, an insulator ultraviolet corona discharge image with a size of 640x640x3 is input into a YOLO-SM detection network for feature extraction. The feature extraction network gradually extracts deep feature representations of the image in a hierarchical progressive manner.

[0034] The overall structure of the feature extraction network includes 5 convolution modules and 12 SCCDG lightweight feature extraction modules. The convolution module, as a basic feature extraction unit, is composed of a two-dimensional convolution layer (Conv2d), a batch normalization layer (BatchNorm) and a SiLU activation function in sequence, with a convolution kernel size of 3x3 and a step size of 2, for realizing feature down-sampling and channel number expansion.

[0035] The input image is first processed by two convolution modules, and the image size is reduced from 640x640 to 160x160, and the channel number is increased from 3 to 128. Subsequently, the following processing is performed in sequence: the first group includes 2 consecutive lightweight feature extraction modules, with an output size of 80x80x256; after one ordinary convolution operation, the second group of 2 consecutive lightweight feature extraction modules is entered, with an output size of 40x40x512; after another ordinary convolution operation, the third group of 4 consecutive lightweight feature extraction modules is entered, with an output size of 20x20x512; after another ordinary convolution, the fourth group of 2 consecutive lightweight feature extraction modules is processed, and the final output size is 20x20x1024.

[0036] The entire feature extraction network includes 10 SCCDG lightweight feature extraction modules. Among them, the output features of the first group, the second group and the third group (with sizes of 80x80x256, 40x40x512 and 20x20x512, respectively) are directly connected to the feature fusion network to provide multi-level feature information.

[0037] The SCCDG lightweight feature extraction module is the core innovation point of the present application, which is specially designed for the corona discharge small target detection task, as shown in Figure 3 The module first separates the input features into n sub-features through an average separation channel operation, with a channel number of C / n for each sub-feature, obtaining n features. The n sub-features are sequentially subjected to convolution operation with a convolution kernel size of 1x1 and a step size of 2, and then the convolved features are stacked in channels to form a unified feature representation.

[0038] The stacked features are input into a CDG module for deep feature extraction, as shown in Figure 4 The CDG module adopts a multi-branch parallel processing architecture and simultaneously performs three different convolution operations on the input features: the first branch adopts a convolution operation with a kernel size of 1x1 and a step size of 1 to obtain features ; the second branch adopts a dilated convolution operation with a kernel size of 3x3 and an expansion rate of 2 to obtain features , which can capture more extensive context information without increasing the amount of calculation by introducing a dilated expansion receptive field; and the third branch adopts a grouped convolution operation with a kernel size of 3x3 and a group number of 2 to obtain features , which significantly reduces the calculation overhead and parameter amount through the grouped convolution.

[0039] The three convolution results are subjected to a channel stacking operation to obtain features , and then are subjected to a two-dimensional convolution operation with a kernel size of 1x1 and a step size of 1, a batch normalization operation, and a Gaussian error activation function operation in sequence to obtain features . Then, the features are subjected to a convolution operation with a kernel size of 1x1 and a step size of 1, an average pooling operation with a kernel size of 3x3 and a padding of 1, and a fully connected layer operation in sequence to obtain features . Finally, the features and are connected in a residual manner to obtain output features .

[0040] The output of the SCCDG lightweight feature extraction module is subjected to a feature integration operation through a last convolution operation with a kernel size of 1x1 and a step size of 1, and finally outputs a multi-scale feature map with a size of HxWxC. Through the combined design of channel separation, multi-branch convolution, and residual connection, this module not only ensures the effective extraction of small target features but also maintains the lightweight calculation characteristics, enabling the network to run efficiently in resource-constrained embedded devices while significantly improving the detection accuracy of small targets such as corona discharge.

[0041] S300: input the multi-scale features into the feature fusion network, generate attention weights through the dual-branch structure of the MSDAM multi-scale dual attention mechanism module, and perform weighted fusion on the multi-scale features to obtain fused features; Furthermore, the MSDAM multi-scale dual attention mechanism module includes: after preprocessing the input features, two parallel branches are input respectively. The first branch is processed by a first activation function after performing average pooling and max pooling operations at the first scale. The second branch is processed by a second activation function after performing average pooling and max pooling operations at the second scale. The two branches form residual connections with the input features to generate attention weights, and finally, weighted fusion is performed.

[0042] Furthermore, the first branch and the second branch use pooling kernels of different sizes for pooling operations. The first branch uses the Mish activation function after pooling, while the second branch uses the Swish activation function after pooling, resulting in different attention weight features between the two branches.

[0043] Specifically, the multi-scale features output from the feature extraction network are input into the feature fusion network for feature integration and optimization. The feature fusion network achieves effective fusion of multi-level features through upsampling, downsampling, and feature concatenation operations, and integrates the MSDAM multi-scale dual attention mechanism module to enhance the feature representation of the target region.

[0044] The overall architecture of the feature fusion network is as follows Figure 2 As shown, the network comprises two upsampling modules, four channel stacking modules, two convolutional modules, four lightweight feature extraction modules, and three multi-scale dual attention mechanism modules. The network adjusts and fuses image features in both spatial scale and channel dimension through upsampling and downsampling operations combined with channel stacking. The three MSDAM multi-scale dual attention mechanism modules connect to three output detection heads at different scales, providing attention-weighted high-quality features for subsequent target prediction.

[0045] The MSDAM multi-scale dual attention mechanism module is a core component of the feature fusion network, such as... Figure 5 As shown, this module is specifically designed to address the problem of indistinct corona discharge target features and susceptibility to background noise interference in complex backgrounds. It employs a dual-branch parallel processing architecture, adaptively enhancing the network's focus on important feature regions through multi-scale information extraction and attention weighting mechanisms.

[0046] The specific implementation process of the MSDAM multi-scale dual attention mechanism module is as follows: The input features first undergo a convolution operation with a kernel size of 1×1 and a stride of 1 to obtain the features. This is used to unify feature dimensions and prepare for subsequent processing. Then, the features... The process proceeds to the two parallel branches, one above the other, for processing.

[0047] In the first branch, features The average pooling operation with a kernel size of 3x3 and padding of 1 and the max pooling operation with a kernel size of 3x3 and padding of 1 are sequentially performed. The average pooling can retain the overall information of the local region, while the max pooling highlights the significant features, and the combination of the two pooling methods effectively extracts different types of spatial information. The results of the two pooling operations are merged through a channel stacking module, and then input into a pointwise convolution module with a kernel size of 1x1 and a step of 1 for feature integration, and then processed by a Mish activation function to obtain feature . The Mish activation function has smooth nonlinear characteristics and can provide better gradient flow and feature expression capability. Feature is constructed into a residual connection with the original input to generate feature with attention weights.

[0048] In the second branch, the processing method is similar to the first branch, but different scale parameters are used. Feature is respectively subjected to average pooling and max pooling operations with a kernel size of 5x5 and padding of 2. Compared with the 3x3 pooling kernel of the first branch, the 5x5 pooling kernel can capture more spatial context information in a larger range. Similarly, the results of the two pooling operations are sent into a pointwise convolution module with a kernel size of 3x3 and padding of 1 after channel stacking, and processed by a Swish activation function to obtain feature . The self-gating property of the Swish activation function can dynamically adjust the information flow and enhance the expression capability of the model. Feature is constructed into another residual connection with the input to output feature with attention weights.

[0049] The two branches can extract feature information from multiple scales and different nonlinear transformations by using different sizes of pooling kernels (3x3 and 5x5) and different activation functions (Mish and Swish), producing attention weight features with complementarity. The first branch focuses on the extraction of local detail features, and the second branch pays more attention to the capture of global context information.

[0050] Finally, the features with attention weights and generated by the two branches are subjected to a weighted fusion operation. Through the learned attention weights, the module can adaptively adjust the importance of different spatial positions and channels in the feature map, highlighting the feature expression of the corona discharge target region while suppressing the interference of background noise. The fused feature is again subjected to a convolution operation with a kernel size of 1x1 and a step of 1 for final feature integration to obtain a fused feature weighted by multi-scale double attention.

[0051] Through the processing of the MSDAM multi-scale dual attention mechanism module, the feature fusion network can effectively reduce the interference of complex background noise on target detection, enhance the recognition ability of the network for the corona discharge target, and significantly improve the detection precision and robustness.

[0052] S400: inputting the fusion features into a prediction network for target detection, and outputting the boundary box coordinates, class label and confidence of the corona discharge target.

[0053] Further, the prediction network comprises three detection heads, each of which comprises a boundary box regression branch, a class classification branch and a confidence prediction branch, and the three detection heads process large-scale, medium-scale and small-scale target detection tasks respectively.

[0054] Further, the output results of the three detection heads are combined for post-processing, including setting a confidence threshold to filter low-confidence detection boxes, removing redundant detection boxes through an IoU threshold by using a non-maximum suppression algorithm, and outputting the final corona discharge target detection result.

[0055] Specifically, the fusion features output by the feature fusion network are input into the prediction network for final target detection. The prediction network is responsible for converting high-level semantic features into specific detection results, and outputs the accurate position, class information and confidence score of the corona discharge target.

[0056] The prediction network adopts a multi-scale detection head architecture, comprising three independent detection heads corresponding to large-scale, medium-scale and small-scale target detection tasks respectively. This multi-scale design can effectively cope with the problem of large size variation of corona discharge targets in actual scenes, ensuring that targets of different sizes can be accurately detected. The three detection heads respectively receive outputs from different levels of MSDAM multi-scale dual attention mechanism modules in the feature fusion network, forming a detection system from coarse granularity to fine granularity.

[0057] Each detection head has the same internal structure and comprises three parallel prediction branches: a boundary box regression branch, a class classification branch and a confidence prediction branch. The boundary box regression branch is responsible for predicting the spatial position information of the target, and outputs four parameters including the boundary box center point coordinates (x, y) and the width w and height h of the boundary box, which realizes accurate positioning of the corona discharge target through regression learning. The class classification branch is used to determine which class the detected target belongs to, and in the application scenario of the present application, it mainly identifies the corona discharge phenomenon, and outputs the probability distribution of each class. The confidence prediction branch evaluates the reliability of the detection result, and outputs a confidence score between 0 and 1, reflecting the probability that there is indeed a target in the detection box.

[0058] The output dimensions of the three detection heads are adapted to different feature map sizes. The large-scale detection head processes smaller feature maps and is mainly responsible for detecting large targets in the image; the medium-scale detection head processes medium-sized feature maps and focuses on medium-sized targets; the small-scale detection head processes larger feature maps and has higher spatial resolution, and is used to detect small target corona discharge phenomena. This hierarchical detection strategy significantly improves the network's detection ability for targets of different scales.

[0059] The original output results of the three detection heads need to go through a post-processing step to get the final detection results. The post-processing process includes two main stages: confidence threshold filtering and non-maximum suppression processing.

[0060] In the confidence threshold filtering stage, the system sets the confidence threshold to 0.5 and filters all the detection boxes output by the three detection heads according to the confidence scores. Only the detection boxes with a confidence score higher than the set threshold are retained for subsequent processing, while the low-confidence detection boxes are discarded directly. This step can effectively remove a large number of false positive detection results and improve the accuracy of detection.

[0061] In the non-maximum suppression (NMS) processing stage, an algorithm based on the IoU (Intersection over Union) threshold is used to remove overlapping redundant detection boxes. The specific process is as follows: first, sort all retained detection boxes in descending order of confidence scores; then select the detection box with the highest confidence score as the reference box; calculate the IoU value of the reference box with all other detection boxes; when the IoU value exceeds the set threshold of 0.5, it is considered that the two detection boxes detect the same target, and the detection box with higher confidence is retained and the detection box with lower confidence is deleted; repeat this process until all detection boxes are processed.

[0062] The final output after post-processing includes the boundary box coordinates of each corona discharge target , the class label "discharge", and the corresponding confidence score. Among them, the boundary box coordinate parameters have the following meanings: : The horizontal coordinate (pixel position) of the upper left corner of the boundary box; : The vertical coordinate (pixel position) of the upper left corner of the boundary box; : The horizontal coordinate (pixel position) of the lower right corner of the boundary box; : The vertical coordinate (pixel position) of the lower right corner of the boundary box.

[0063] These four coordinate values determine the exact position and range of the corona discharge target in the original image and can be directly used to draw detection boxes on the ultraviolet image to realize the visualization labeling of corona discharge phenomena, providing accurate and reliable technical support for the safety monitoring and fault diagnosis of power transmission lines. The design of the entire prediction network ensures high precision and high reliability of the detection results, meeting the needs of practical engineering applications.

[0064] Embodiment Two A power company applied the method to detect UV corona discharge of insulators on a 220 kV transmission line. The transmission line is about 50 kilometers long, with 120 towers, each equipped with 12-15 strings of composite insulators. Due to the serious salt pollution in the coastal area, corona discharge often occurs on the surface of the insulators, which needs to be regularly inspected and monitored.

[0065] The traditional detection method has the following problems: low efficiency of manual inspection, 15 days of full-line inspection each time; low accuracy of night UV imaging detection, with a missed detection rate of 25%; existing target detection algorithms have difficulty in identifying small-size corona discharge targets, with a false alarm rate of up to 30%.

[0066] The company uses the insulator UV corona discharge target detection method based on YOLO-SM to deploy the detection system on the UAV-mounted UV imaging device. The system configuration is as follows: mounted with Jetson AGXXavier edge computing device, 32GB memory, 512GB SSD storage; UV camera resolution 1024x768, frame rate 30fps; detection distance 50-200 meters.

[0067] In a one-month practical application, the method was used to continuously monitor the transmission line, and 5000 UV images were obtained, including 800 images containing corona discharge targets. Compared with traditional YOLOv5, YOLOv8 and manual detection methods, the following experimental results were obtained:

[0068] Table 1. Test comparison table

[0069] As shown in Table 1, through comparative analysis, it can be seen that the method is significantly better than the existing methods in various indicators: the detection accuracy of the method reaches 92.4%, which is 7.1 percentage points higher than YOLOv8 and 13.9 percentage points higher than manual detection. This is mainly due to the SCCDG lightweight feature extraction module, which can better extract the subtle features of small target corona discharge through multi-branch convolution processing.

[0070] Recall rate is greatly improved: the recall rate reaches 89.7%, the missed detection rate is reduced to 10.3%, which is 17.4 percentage points lower than the traditional method. The MSDAM multi-scale dual attention mechanism module effectively enhances the recognition ability of weak corona signals in complex background through different scale pooling and dual activation function.

[0071] False alarm rate is significantly reduced: the false alarm rate is only 7.6%, which is 7.1 percentage points lower than YOLOv8. The shape deviation loss and corner distance loss terms in the composite bounding box regression loss function enable the model to more accurately locate irregularly shaped corona discharge areas, reducing false positives from background noise.

[0072] Computational efficiency is improved: the detection speed reaches 32.5ms / frame, which is 16% higher than YOLOv8, and the model size is only 18.9MB, making it easy to deploy on edge devices. The channel separation and grouped convolution design in the SCCDG lightweight feature extraction module effectively reduces computational overhead.

[0073] Actual application effect: In 1 month of actual operation, the system successfully detected 43 corona discharge fault points, of which 35 were confirmed as real faults by the field, with an accuracy of 81.4%. Compared with the previous manual inspection of 20-30 fault points per month, the detection efficiency has improved significantly. The system can also achieve 7x24 continuous monitoring, timely detecting corona discharge phenomena at night and in adverse weather conditions, providing a strong guarantee for the safe and stable operation of the power transmission line.

[0074] In a certain power laboratory, the effectiveness of the core technology components of the application was verified, and actual 500kV substation insulator ultraviolet corona discharge images were used for comparative testing. The test data were 600 ultraviolet images collected on site at a 500kV substation.

[0075] The performance differences between using a standard convolution module and the SCCDG lightweight feature extraction module were compared: Table 2. SCCDG lightweight feature extraction module effectiveness verification table

[0076] As shown in Table 2, the test results show that the SCCDG lightweight feature extraction module, through channel separation and multi-branch convolution design, significantly reduces model complexity while improving detection accuracy.

[0077] The detection effects with and without the MSDAM multi-scale dual attention mechanism were compared: Table 3. MSDAM multi-scale dual attention mechanism module attention mechanism verification table

[0078] As shown in Table 3, the MSDAM multi-scale dual attention mechanism module, through the cooperation of double-branch multi-scale pooling and different activation functions, effectively enhances the recognition ability of small target corona discharge in complex background.

[0079] The effects of traditional intersection over union loss and composite bounding box regression loss function were compared: Table 4. Composite loss function verification table

[0080] As shown in Table 4, the composite loss function significantly improves the detection and positioning accuracy of irregularly shaped corona discharge targets by fusing shape deviation loss, corner distance loss, and density weighted loss.

[0081] In the actual test of a 500kV substation, the complete YOLO-SM detection system showed obvious performance advantages compared to the version with each core component removed: Overall detection accuracy: improved from 84.7% for the basic version to 92.4% for the complete version.

[0082] Computing efficiency: model size reduced by 25.3%, inference speed increased by 21.1%.

[0083] Practicality: false detection rate in complex power environments reduced to 7.6%, meeting the needs of actual engineering applications.

[0084] This test confirms the independent contribution and synergistic optimization effect of each core technology component of the present application, providing a reliable technical solution for power system insulator ultraviolet corona discharge detection.

[0085] Finally, it should be noted that: the above only for the preferred embodiments of the present application, and not for the limitation of the present application, although the foregoing embodiments of the present invention are described in detail, for those skilled in the art, it still can be modified, or part of the technical features of the equivalent replacement. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for detecting target of ultraviolet corona discharge of insulator based on YOLO-SM, characterized in that, The method comprises the steps of: S100: obtaining a trained YOLO-SM detection network, which comprises a feature extraction network, a feature fusion network and a prediction network, wherein: The feature extraction network is integrated with an SCCDG lightweight feature extraction module, which separates multiple sub-features through an average channel separation operation and performs parallel ordinary convolution, dilated convolution and grouped convolution on the input features through a CDG module; The feature fusion network is integrated with an MSDAM multi-scale dual attention mechanism module, which comprises two parallel branches that use different size pooling kernels and are processed by Mish and Swish activation functions respectively to generate complementary attention weights; The YOLO-SM detection network is trained using a composite bounding box regression loss function; S200: inputting the insulator ultraviolet corona discharge image into the YOLO-SM detection network, and performing channel splitting and multi-branch convolution processing on the input features through the SCCDG lightweight feature extraction module integrated in the feature extraction network to extract multi-scale features; S300: inputting the multi-scale features into the feature fusion network, performing different scale pooling operations through the double-branch structure of the MSDAM multi-scale dual attention mechanism module to generate attention weights, and weighting and fusing the multi-scale features to obtain fused features; S400: inputting the fused features into the prediction network for target detection, and outputting the bounding box coordinates, class label and confidence of the corona discharge target.

2. The YOLO-SM-based insulator ultraviolet corona discharge target detection method according to claim 1, characterized in that, The composite bounding box regression loss function includes the weighted sum of the following four losses: The intersection over union loss is calculated by the loss value of the overlap between the predicted box and the target box; The shape deviation loss is calculated by the difference in aspect ratio between the predicted box and the target box; The corner distance loss is calculated by the distance between the corner coordinates of the predicted box and the target box; The density weighted loss sets a weighting factor based on the target distribution density.

3. The YOLO-SM-based insulator ultraviolet corona discharge target detection method according to claim 1, characterized in that, The feature extraction network comprises a convolution module and an SCCDG lightweight feature extraction module, and the convolution module is composed of two-dimensional convolution, batch normalization and activation function, and the outputs of multiple SCCDG lightweight feature extraction modules are connected to the feature fusion network to provide multi-level features.

4. The YOLO-SM-based insulator ultraviolet corona discharge target detection method according to claim 1, characterized in that, The SCCDG lightweight feature extraction module comprises: The input features are separated into multiple sub-features through an average channel separation operation, the sub-features are stacked after convolution operation, and the stacked features are input into the CDG module for processing, the CDG module outputs are connected to the stacked features to form a residual connection, and finally the processed features are output through convolution operation; The CDG module performs three-way convolution processing on the input features in parallel: the first way performs ordinary convolution operation, the second way performs dilated convolution operation, and the third way performs grouped convolution operation; the three convolution results are stacked in the channel, and then processed through convolution, normalization and activation function, and finally connected to the intermediate processing results through convolution and pooling operations.

5. The YOLO-SM-based insulator ultraviolet corona discharge target detection method according to claim 1, characterized in that, The feature fusion network comprises an up-sampling module, a channel superposition module, a convolution module, an SCCDG lightweight feature extraction module and an MSDAM multi-scale dual attention mechanism module, the feature scale is adjusted through up-sampling and down-sampling operations, feature splicing is realized through channel superposition, and the multiple MSDAM multi-scale dual attention mechanism modules are respectively connected with detection heads of different scales.

6. The YOLO-SM-based insulator ultraviolet corona discharge target detection method according to claim 1, characterized in that, The MSDAM multi-scale dual attention mechanism module comprises: input features are respectively input into two parallel branches after preprocessing, the first branch is processed through a first activation function after average pooling and maximum pooling operations of a first scale, the second branch is processed through a second activation function after average pooling and maximum pooling operations of a second scale, the two branches respectively form residual connections with input features to generate attention weights, and finally weighted fusion is performed.

7. The YOLO-SM-based insulator ultraviolet corona discharge target detection method according to claim 6, characterized in that, The first branch and the second branch perform pooling operations through different sizes of pooling kernels, the first branch is processed through a Mish activation function after pooling, the second branch is processed through a Swish activation function after pooling, and the two branches generate different attention weight features.

8. The YOLO-SM-based insulator ultraviolet corona discharge target detection method according to claim 1, characterized in that, The prediction network comprises three detection heads, each detection head comprises a bounding box regression branch, a class classification branch and a confidence prediction branch, and the three detection heads respectively process large-scale, medium-scale and small-scale target detection tasks.

9. The YOLO-SM-based insulator ultraviolet corona discharge target detection method according to claim 8, characterized in that, The output results of the three detection heads are combined and post-processed, including setting a confidence threshold to filter low-confidence detection boxes, removing redundant overlapping detection boxes through an IoU threshold by using a non-maximum suppression algorithm, and outputting final corona discharge target detection results.

Citation Information

Patent Citations

  • Function-level code vulnerability detection method based on slice attribute graph representation learning

    CN112699377A

  • Safety vest target detection method based on YOLOv7 algorithm

    CN120259705A

  • Remote sensing target detection method, equipment and medium

    CN120431479A

  • Ship sign detection method based on multi-scale feature fusion and attention mechanism

    CN120431564A

  • Multi-modal image ship target detection method and system

    CN120599231A