Power transmission line icing detection method based on PCTE-Net

By using the improved PCTE-Net model, which employs point-to-point connected serpentine convolution, target-enhanced grouped global normalized attention, and lightweight double-shared normalized convolution, the environmental and geographical limitations in icing detection are addressed, achieving high-precision icing recognition.

CN121527518AInactive Publication Date: 2026-02-13ELECTRIC POWER SCI RES INST OF STATE GRID XINJIANG ELECTRIC POWER CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511701147.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-02-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies for icing detection are limited by natural environment and geographical conditions, resulting in large detection errors and making it difficult to accurately identify the icing situation of transmission lines.

Method used

An improved PCTE-Net model is adopted, which combines pointwise connected serpentine convolution DCSConv, target-enhanced grouped global normalized attention module EGNA, and lightweight double-shared normalized convolution LWS-NormConv to improve the model's ability to detect ice in complex environments.

Benefits of technology

It can accurately identify icing on transmission lines in harsh environments, improving detection accuracy and robustness, with an mAP of 88.1%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121527518A_ABST
    Figure CN121527518A_ABST
Patent Text Reader

Abstract

The invention discloses a PCTE-Net-based power transmission line icing detection method. The method comprises the following steps: S1, obtaining an icing image of a power transmission line; s2, dividing the power transmission line icing data set into a training set, a verification set and a test set; s3, adopting a data labeling tool to label icing types; s4, training the training set marked in the step S3 by adopting an improved PCTE-Net model; s5, performing detection by using the model trained in the step S4; and S6, evaluating the detection result in the S5 by adopting the accuracy rate, the recall rate and the mAP value, the method constructs a detection network for effectively identifying icing in the power transmission line under complex conditions, the improved PCTE-Net model firstly adopts point-by-point connected domain snaking convolution, and then proposes a target enhanced grouping global standardized attention module EGNA, and the detection result in the step S5 is evaluated by adopting the accuracy rate, the recall rate and the mAP value. The detection precision and robustness are improved while the light weight is kept by using the light-weight dual-use normalized convolution, and finally, the icing area of the power transmission line is accurately positioned in the severe environment, and the icing type is accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing for power maintenance, and in particular to a transmission line icing detection method based on PCTE-Net. BACKGROUND

[0002] With the rapid development of social economy, the demand for electricity continues to rise. In China, the long-term high growth of electricity consumption has driven the implementation of new power generation projects, especially in the construction of ultra-high voltage, super-high voltage transmission and long-distance overhead transmission lines. At the same time, the power system is also accelerating towards intelligentization and automation. However, as the scale of the power grid continues to expand, its operation safety has become a growing concern. As the core link of power transmission, the safety of the transmission line is crucial to the protection of the national economy and social life. China is vast in territory and has diverse climates, and is one of the countries most severely affected by icing disasters in the world. Icing refers to the continuous icing and accumulation on the surface of the transmission line under the conditions of heavy snow, freezing rain or sustained low temperature, which leads to increased conductor sag, increased insulator stress, increased line load and possible twisting phenomena. In severe cases, icing can cause conductor breakage, hardware damage, tower tilting and even collapse, which seriously threatens the safe and stable operation of the power system.

[0003] In recent years, China's power system has been severely affected by icing disasters several times. In January 2018, affected by the East Asian cold wave, serious icing phenomena occurred again in central and southwestern China, and some transmission lines experienced galloping and failures. The local power department organized large-scale emergency repair and restored power supply in time. On the evening of December 13, 2023, affected by the snow and cold weather, 4 transmission lines in Yuanqu County, Shanxi Province failed, 12 substations in Shanxi were shut down, 55,000 people, 547 power generation vehicles and 2,862 power generators were used to carry out power repair work. In early 2024, the mountainous area of Xiangxi suddenly encountered freezing rain, and the icing thickness of the 220-kilovolt transmission line exceeded the design standard. The power company quickly carried out artificial deicing and equipment maintenance, effectively avoiding regional power outages. On the eve of the 2025 Spring Festival, the icing of the transmission line in the Qinling Mountains caused the monitoring device to malfunction, and the State Grid operation and maintenance personnel completed the equipment replacement and debugging in the harsh environment, ensuring the safety of railway transportation and residential electricity. The above cases show that icing disasters have become a key factor threatening the safe operation of China's power grid, and it is of great practical significance to strengthen the research on icing monitoring and prevention technology.

[0004] Therefore, in order to timely discover and solve the problem of icing on the transmission line, and improve the safety and stability of the transmission line, by monitoring the state of the transmission line in real time, the icing condition of the transmission line can be discovered in time, and appropriate maintenance measures can be taken to ensure the safe operation of the transmission line. SUMMARY

[0005] In order to overcome the above problems, the purpose of the present application is to provide a power line icing detection method based on PCTE-Net, which solves the shortcomings of the prior art, such as being limited by natural environment and geographical conditions, and detection errors caused by the characteristics of icing itself.

[0006] The technical scheme adopted by the present application is:

[0007] The power line icing detection method based on PCTE-Net comprises the following steps:

[0008] S1: Obtain the power line icing image.

[0009] S2: Divide the power line icing data set obtained in S1 into a training set, a validation set and a test set, and the ratio of the training set, the validation set and the test set is 8:1:1.

[0010] S3: Use a data labeling tool to label the data in the training set in S2 according to the icing class, and obtain a power line icing database.

[0011] S4: Train the training set labeled in S3 using the improved PCTE-Net model.

[0012] S5: Use the trained model in S4 to detect the data in the validation set in S2, and obtain the corresponding power line icing detection result.

[0013] S6: Evaluate the detection result in S5 using precision, recall and mAP value, and obtain the efficiency of power line icing detection.

[0014] As a further description of the present application, the improved PCTE-Net model in S4 is improved based on the YOLOv11 model, and the improved part includes:

[0015] S41: Replace the standard 3x3 convolution in the C3K2 module of the main stem with the point-by-point connected domain snake convolution DCSConv.

[0016] S42: Add the target enhanced grouping global normalization attention module EGNA to the part after the basic feature extraction of each scale feature through C3k2.

[0017] S43: Use the lightweight double sharing normalization convolution LWS-NormConv to improve the original detection head in the head.

[0018] As a further description of the present application, the specific process in S1 is:

[0019] S11: Use a UAV to patrol and take pictures of the icing image on the power line.

[0020]

[0020] S12: Expand the dataset by performing horizontal flipping, random cropping, rotation transformation, brightness enhancement, and contrast enhancement operations on the obtained icing images of transmission lines to obtain an expanded dataset.

[0021] S13: Reintegrate the icing images of transmission lines captured in S11 with the expanded icing images of transmission lines into a new dataset as a sample library.

[0022] As a further description of the present invention, the specific process of S3 is as follows:

[0023] S31: Select the icing images of the transmission lines to be labeled, and divide the icing images of the transmission lines into four labels: snow, rime, mixed rime, and ice.

[0024] S32: Use the Labellmg annotation software to generate XML tag files corresponding to the icing images of transmission lines, and store the annotated sample data in Pascal VOC format.

[0025] As a further description of the present invention, the specific process of point-by-point connected domain serpentine convolution DCSConv in S41 is as follows:

[0026] S411: Let K be the coordinate set of the standard two-dimensional convolution kernel, and let its center point be... For a 3×3 convolution kernel, the coordinates can be represented as:

[0027] .

[0028] S412: Apply the standard convolution kernel to... shaft and The axial directions are linearized respectively.

[0029] exist On the axis, a convolution kernel of size 9, its th... Each sampling point is represented as:

[0030] , .

[0031] Each position is generated iteratively from the center point in sequence:

[0032] .

[0033] .

[0034] exist Similarly, in the axial direction, it is defined as:

[0035] .

[0036] .

[0037] in, They are respectively shaft and The offset of the axis. for The center point along the axial direction for The center point along the axis.

[0038] S413: The final sampling points are obtained using bilinear interpolation. The formula for calculating the final sampling points using bilinear interpolation is as follows:

[0039] ,

[0040] in, For the nearest integer coordinates, This is the interpolation kernel function.

[0041] As a further description of the present invention, the specific process of the Target Enhanced Grouped Global Normalized Attention Module (EGNA) in S42 is as follows:

[0042] S421: Divide the input feature map into several independent groups according to channels. Each group contains a portion of channels and processes them independently.

[0043] S422: Add and merge all channel features within each group to extract the most critical spatial information for that group.

[0044] S423: The Sigmoid function is used to compress the result to between 0 and 1 to generate an attention mask.

[0045] S424: Multiply the generated attention mask with the feature correspondences of the initial group.

[0046] S425: Re-merge all grouped enhanced features to restore the feature size to the input size, and obtain the final enhanced result.

[0047] As a further description of the present invention, the specific process of the lightweight double-shared normalized convolution LWS-NormConv in S43 is as follows:

[0048] S431: Introducing a batch normalization mechanism, each input feature (P3, P4, P5) goes through the BN_Conv 1×1 module to perform linear transformation and normalization on the feature channels.

[0049] S432: Introducing kernel weight sharing to reduce parameter quantity and redundancy. In the core part of the detection head, the two BN_Conv 3×3 modules adopt a kernel weight sharing mechanism, that is, features of different scales use the same weights in the convolution operation.

[0050] S433: Introduce a shared strategy in the detection branch. In the prediction phase, the detection head of the detection branch is divided into three parts: bounding box regression branch (Conv_Box+Scale), category classification branch (Conv_Cls), and mask prediction branch (Conv_Mask).

[0051] All Conv_Box modules share a set of convolutional weights, and the Scale module learns independent scaling factors for each scale.

[0052] All Conv_Cls modules share the same set of convolution weights.

[0053] All Conv_Mask modules share the same convolution kernel parameters.

[0054] As a further description of the present invention, the formula for calculating the accuracy in S6 is as follows:

[0055] .

[0056] Where P represents precision. For the detected positive examples, These are the negative examples detected.

[0057] The formula for calculating the recall rate is:

[0058] .

[0059] In the formula, R is the recall rate. For the detected positive examples, These are positive examples that were not detected.

[0060] The formula for calculating the mAP value is:

[0061] .

[0062] in, The number of target categories involved in the calculation. Let R be the precision function under the condition of recall. Let R be the derivative with respect to recall rate R.

[0063] As a further description of the present invention, S6 can also be evaluated using frames per second, parameter quantity, and number of floating-point operations.

[0064] The beneficial effects of this invention are:

[0065] This invention presents a PCTE-Net-based method for detecting icing on transmission lines. The method constructs a detection network capable of effectively identifying icing on transmission lines under complex conditions. First, it employs point-to-point connected serpentine convolutions, maintaining the continuity of the convolution kernel's receptive field and its fit to slender structures while preserving a certain degree of offset freedom, allowing the model to flexibly adapt to complex morphologies. Second, it proposes a target-enhanced grouped global normalized attention module (EGNA), utilizing a channel grouping strategy to differentiate and enhance the features of different icing morphologies, improving the model's adaptability to diverse icing conditions. Furthermore, through dynamic learning of spatial attention weights, it adaptively focuses on low-contrast, small-scale icing regions. Finally, it uses lightweight double-shared normalized convolutions (LWS-NormConv) to improve detection accuracy and robustness while maintaining lightweight design. Experimental data confirms that the network designed in this invention can accurately locate icing areas on transmission lines in harsh environments, precisely identifying icing, rime, snow, and mixed rime types, with an mAP of 88.1%. Attached Figure Description

[0066] Figure 1 This is the overall flowchart of the present invention.

[0067] Figure 2 This is a network structure diagram of the PECT-Net algorithm of the present invention.

[0068] Figure 3 This is a diagram of the point-by-point connected domain serpentine convolution structure of the present invention.

[0069] Figure 4 This is a structural diagram of the target-enhanced attention module of the present invention.

[0070] Figure 5 This is a diagram of the lightweight double-shared normalized convolution structure of the present invention.

[0071] Figure 6 This is a comparison chart of the thermal detection results of the present invention. Detailed Implementation

[0072] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0073] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0074] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0075] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.

[0076] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0077] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0078] like Figures 1-6 As shown, it illustrates a specific embodiment of the present invention:

[0079] Example 1:

[0080] A method for detecting icing on transmission lines based on PCTE-Net, such as Figure 1 As shown, it includes the following steps:

[0081] S1: Acquisition of images of icing on transmission lines.

[0082] In this embodiment, images of icing on power transmission lines in different environments are obtained through drone inspections. These images are then augmented, and the augmented images are used as a dataset of real icing on power transmission lines.

[0083] S2: Divide the transmission line icing dataset obtained in S1 into a training set, a validation set, and a test set, with the ratio of the training set, validation set, and test set being 8:1:1.

[0084] S3: Use data labeling tools to label the icing categories of the training set data in S2 to obtain the transmission line icing database.

[0085] In this embodiment, the Labellmg software is used for annotation, and the image annotation is performed in VOC format. The icing image information of the transmission line after annotation is saved to obtain the transmission line icing database.

[0086] S4: Train the training set labeled in S3 using the improved PCTE-Net model.

[0087] S5: Use the model trained in S4 to detect the data in the validation set in S2, and obtain the corresponding transmission line icing detection results.

[0088] S6: The detection results in S5 are evaluated using precision, recall, and mAP value to determine the efficiency of transmission line icing detection.

[0089] Example 2:

[0090] Specifically, the improved PCTE-Net model structure used in S4 is as follows: Figure 2 As shown, this model is an improvement upon the YOLOv11 model, and the improvements include:

[0091] S41: The standard 3×3 convolution in the C3K2 module of the backbone is replaced by DCSConv using point-by-point connected serpentine convolution. The point-by-point connected serpentine convolution maintains the continuity of the receptive field of the convolution kernel and the fit to slender structures, while also retaining a certain degree of offset freedom, so that the model can flexibly adapt to complex shapes.

[0092] In this embodiment, the point-by-point connected serpentine convolution borrows the concept of a "serpent," causing the convolution kernel to crawl continuously along the target curve point by point, like a controlled snake. Figure 3 As shown, this is the specific structure of the serpentine convolution. This "point-by-point cumulative offset" mechanism maintains the continuity of the convolution kernel's receptive field and its fit to slender structures, while also preserving a certain degree of offset freedom, allowing the model to flexibly adapt to complex shapes, such as... Figure 2As shown, the specific process of point-by-point connected domain serpentine convolution DCSConv is as follows:

[0093] S411: Let K be the coordinate set of the standard two-dimensional convolution kernel, and let its center point be... For a 3×3 convolution kernel, the coordinates can be represented as:

[0094] .

[0095] S412: Apply the standard convolution kernel to... shaft and The axial directions are linearized respectively.

[0096] exist On the axis, a convolution kernel of size 9, its th... Each sampling point is represented as:

[0097] , .

[0098] Each position is generated iteratively from the center point in sequence:

[0099] .

[0100] .

[0101] exist Similarly, in the axial direction, it is defined as:

[0102] .

[0103] .

[0104] in, They are respectively shaft and The offset of the axis. for The center point along the axial direction for Center point along the axis;

[0105] In this embodiment, to give the convolution kernel greater flexibility and enable it to focus on the complex geometric features of the target, an offset is introduced on the basis of traditional convolution. However, if the offset is learned entirely by the model, the receptive field often deviates from the target, especially when dealing with slender tubular structures. To address this, DCSConv employs an iterative update strategy: gradually determining the next sampling position to ensure the spatial continuity of the region of interest and avoid excessive diffusion of the receptive field due to excessive offset.

[0106] S413: The final sampling points are obtained using bilinear interpolation. The formula for calculating the final sampling points using bilinear interpolation is as follows:

[0107] ,

[0108] in, For the nearest integer coordinates, This is the interpolation kernel function.

[0109] In this embodiment, due to the offset Since the values ​​are usually decimals, and image coordinates are discrete integer grids, bilinear interpolation is needed to obtain the final sampling points. The interpolation kernel function is decomposed into two one-dimensional kernels, one horizontally and one vertically, and the calculation formula is as follows:

[0110] .

[0111] in, Let K be the coordinates of the continuous sampling point K along the x-axis. For its nearest integer grid points The coordinates in the x-axis direction, Let K be the coordinates of the continuous sampling point K along the y-axis. Neighboring integer grid points The coordinates in the y-axis direction, It is a one-dimensional linear interpolation kernel function.

[0112] S42: The Target Enhancement Grouped Global Normalization Attention Module (EGNA) is added to the part after the basic feature extraction of each scale feature by C3k2. In this embodiment, the Target Enhancement Grouped Global Normalization Attention Module (EGNA) uses a channel grouping strategy to achieve differentiated enhancement of icing features of different morphologies. It performs "secondary refinement" of the feature map through a grouped spatial attention mechanism to improve the model's adaptability to icing diversity. On the other hand, through dynamic learning of spatial attention weights, it adaptively focuses on low-contrast small-scale icing areas.

[0113] In this embodiment, after basic feature extraction at each scale using C3k2, the feature map is "refined" a second time through a target-enhanced attention module. On one hand, a channel grouping strategy is used to differentiate and enhance features of different icing morphologies, improving the model's adaptability to icing diversity. On the other hand, through dynamic learning of spatial attention weights, the model adaptively focuses on low-contrast, small-scale icing regions. The specific structure of the target-enhanced grouping global normalization attention module is as follows... Figure 4 As shown, the specific process is as follows:

[0114] S421: Divide the input feature map into several independent groups according to channels. Each group contains a portion of channels and processes them independently.

[0115] In this embodiment, the method allows the model to learn different types of features separately, reducing interference between different features. For each group of features, the global average information is calculated first, and then this global information is multiplied by the original features, so that the features at each position are integrated into the global context.

[0116] S422: Add and merge all channel features within each group to extract the most critical spatial information for that group.

[0117] In this embodiment, to make the feature differences between different spatial locations more obvious, the merged features are standardized: first, the mean of the entire space is subtracted, and then divided by the standard deviation. This step amplifies the difference between the target region and the background, making it easier for the subsequent attention mechanism to distinguish important regions. The standardized features are then restored to their original spatial dimensions, and the importance of each group is adjusted using two sets of learnable parameters.

[0118] S423: The Sigmoid function is used to compress the result to between 0 and 1 to generate an attention mask.

[0119] In this embodiment, the larger the value of the attention mask at each position, the more important the feature at that position is.

[0120] S424: Multiply the generated attention mask with the feature correspondences of the initial group.

[0121] In this embodiment, this step amplifies the features of important regions and suppresses unimportant regions.

[0122] S425: Re-merge all grouped enhanced features to restore the feature size to the input size, and obtain the final enhanced result.

[0123] In this embodiment, let the input feature map be... ,in For batch size, For the number of channels, For spatial dimensions, the goal of the EGNA module is to enhance key features using grouping strategies and spatial attention mechanisms.

[0124] Arrange the channel dimensions into a preset number of groups. Division (must meet) Each group contains Each channel, the input feature map is reshaped as follows:

[0125] .

[0126] in, This is the grouped feature tensor after being re-divided and rearranged according to the number of groups g. The input is the original feature map. The spatial dimension representation of the tensor. This is a tensor shape transformation operation used to rearrange the channel dimensions in groups, so that each group can be processed independently and avoid interference from different channel features.

[0127] To capture global information for each group of features, global average pooling is applied to the grouped features:

[0128] ,

[0129] in, This refers to the group-level global information obtained after performing global average pooling on each group of features. This is a global average pooling operation used to compress each set of features into a 1×1 statistic.

[0130] By fusing global semantic vectors with local features, intra-group information is aggregated along the channel dimension:

[0131] ,

[0132] in, For the first in the group Channel characteristics, The global average pooling feature corresponding to the k-th group. The aggregated features are used to achieve local and global interaction through dot product operations, highlighting key spatial regions.

[0133] To enhance the distinguishability of different spatial locations, the aggregated features are... Perform mean-standard deviation normalization.

[0134] First, flatten out the spatial dimensions:

[0135] .

[0136] in, To aggregate features Reconstruct the feature tensor after expanding the channel group.

[0137] Subtract the mean and divide by the standard deviation:

[0138] .

[0139] .

[0140] .

[0141] .

[0142] in, The centering result after removing the mean of each feature group, Let t be the average value of feature t across each spatial dimension. This is the result after normalizing the centered features according to the standard deviation. Let be the standard deviation of feature t in each spatial dimension. To prevent division by zero of small constants.

[0143] The standardized features are restored to their spatial shape, and learnable parameters are introduced through affine transformation:

[0144] .

[0145] in, For normalized features The group weight features obtained after applying an affine transformation These are learnable weights and biases, used to dynamically adjust the importance of different groups.

[0146] The weights are mapped to using the Sigmoid activation function. We obtain the spatial attention mask:

[0147] .

[0148] in, For spatial attention mask, The Sigmoid activation function is used to map weights to the [0,1] interval, generating a spatial attention mask.

[0149] Applying it to the original grouping features to achieve spatial region enhancement:

[0150] .

[0151] in, This represents element-wise multiplication. This refers to the enhanced features.

[0152] Finally, the enhanced features are restored to the input dimension:

[0153] .

[0154] The final output feature map is obtained. .

[0155] S43: The original detection head in the head is improved by using lightweight double-shared normalized convolution LWS-NormConv. In this embodiment, lightweight double-shared normalized convolution LWS-NormConv is used to improve detection accuracy and robustness while maintaining lightweight design.

[0156] In this embodiment, the original segmentation detection head is improved by introducing Batch Normalization (BN) and convolution weight sharing mechanisms, thereby improving detection accuracy and robustness while maintaining lightweight design. Traditional detection heads typically use a conventional stacking of BatchNorm + Conv modules in their branching structure. Each branch independently constructs the detection head at different scales (P3, P4, P5), resulting in a large number of parameters and potential computational redundancy during inference. In LWS-NormConv, we further introduce a combined strategy of BN and convolution weight sharing on top of the lightweight detection head to effectively reduce computational overhead while maintaining expressive power. The specific structure is as follows: Figure 5 As shown, the specific process of Lightweight Double-Shared Normalized Convolution LWS-NormConv is as follows:

[0157] S431: Introducing a batch normalization mechanism, each input feature (P3, P4, P5) goes through the BN_Conv 1×1 module to perform linear transformation and normalization on the feature channels.

[0158] In this embodiment, Batch Normalization (BN) alleviates the internal covariate shift problem at this stage, ensuring the feature distribution remains stable across different mini-batches and improving the effectiveness of gradient propagation. Compared to the original detector head, this normalization method strengthens the consistency of multi-scale features during feature preprocessing, thus providing a more stable input for subsequent convolution sharing. Secondly, BN parameters are not shared, adapting to multi-scale feature differences. Although the convolution kernels are shared across multiple scales, the parameters of each BN layer remain independent, allowing feature maps at different scales to learn normalization strategies based on their own distribution characteristics. This mechanism is more parameter-efficient than the independent convolution design of the original detector head, while avoiding the normalization failure problem caused by scale differences.

[0159] S432: Introducing kernel weight sharing to reduce parameter quantity and redundancy. In the core part of the detection head, the two BN_Conv 3×3 modules adopt a kernel weight sharing mechanism, that is, features of different scales use the same weights in the convolution operation.

[0160] In this embodiment, this design significantly reduces the number of model parameters. Compared with the original independent convolutional structure, it can reduce computational redundancy and speed up inference without sacrificing feature extraction capabilities.

[0161] S433: Introducing a shared strategy in the detection branch, during the prediction phase, the detection head of the detection branch is divided into three parts: bounding box regression branch (Conv_Box+Scale), category classification branch (Conv_Cls), and mask prediction branch (Conv_Mask).

[0162] Specifically, all Conv_Box modules share a set of convolutional weights, and the Scale module learns independent scaling factors for each scale; all Conv_Cls modules share the same set of convolutional weights; and all Conv_Mask modules share convolutional kernel parameters.

[0163] In this embodiment, during the prediction phase, the detection head of the detection branch is still divided into three parts: bounding box regression branch (Conv_Box+Scale), category classification branch (Conv_Cls), and mask prediction branch (Conv_Mask). All Conv_Box modules share a set of convolutional weights to maintain consistency in bounding box regression across different scales. Simultaneously, the Scale module learns independent scaling factors for each scale, thereby enhancing adaptability to multi-scale targets. All Conv_Cls modules share the same set of convolutional weights, ensuring consistent feature representation for category prediction across different scales. All Conv_Mask modules share convolutional kernel parameters to generate a consistent mask representation on multi-scale feature maps. Through this design of shared convolutional weights and independent normalized parameters, LWS-NormConv effectively reduces the number of detection head parameters while ensuring the collaborative expressive ability of classification, localization, and segmentation tasks across different scales, further improving overall detection and segmentation performance.

[0164] Compared to existing segmentation and detection heads, LWS-NormConv maintains a lightweight network while enhancing the stability of feature distribution through the normalization advantage of Batch Normalization (BN). It adapts to the variability of multi-scale features by using non-shared BN parameters and reduces overall redundancy by sharing convolutional kernels. Ultimately, this structure achieves more robust performance in both small and large target detection, balancing accuracy, efficiency, and generalization ability.

[0165] In this embodiment, the input feature map is assumed to be: Where B is the batch size, C is the number of channels, and H and W are the spatial dimensions.

[0166] In LWS-NormConv, each scale feature First, after passing through the BN_Conv 1×1 module, the formula for calculating Batch Normalization (BN) is:

[0167] .

[0168] .

[0169] in, This is the channel mean; Channel variance; These are learnable scaling and translation parameters; The input feature is the original value of the b-th sample, c-th channel, at spatial location (h, w). The input features are normalized by mean centering and divided by the channel standard deviation. To apply channel scaling to normalized features With translation The resulting BN output, is the standard deviation of the input feature over channel c, used to measure the distribution scale of the channel feature.

[0170] The key is that the BN parameters are independent at each scale in LWS-NormConv:

[0171] .

[0172] This allows the normalization strategy to be more flexible in adapting to multi-scale features.

[0173] After BN_Conv1×1 processing, the feature map is fed into the shared convolutional kernel. The convolution calculation process is as follows:

[0174] ,

[0175] in, This indicates the number of lines processed by BN_Conv1×1. Each scale; features across all scales share the same convolutional kernel. Convolutional weights are shared, but BN parameters are not.

[0176] In the LWS-NormConv detector head, the output features are processed by a shared convolutional module and then divided into three parallel branches: the Box Regression Branch, the Classification Branch, and the Segmentation / Mask Branch. Each branch is designed to be lightweight (using a shared convolutional kernel) while also being independently parameterized through Batch Normalization (BN) to adapt to feature distributions at different scales.

[0177] S434: Obtain the final comprehensive loss function. The final optimization goal of LWS-NormConv is:

[0178] ,

[0179] in, For classification loss (CE / BCE); The regression loss is (CIoU / GIoU). For segmentation mask loss (BCE / Dice Loss); This is the loss weighting coefficient.

[0180] Specifically, the bounding box regression branch is responsible for predicting the target's location information, and the calculation formula is as follows:

[0181] .

[0182] in, Let i be the unscaled position offset predicted by the bounding box regression branch at scale i. To share the convolution function, Let i be the feature map input to the bounding box regression branch at scale i. For shared bounding box convolution kernels.

[0183] Adjusting regression results at different scales using the Scale module:

[0184] .

[0185] in, For scale The learnable scaling factor, This represents the final bounding box regression result at scale i.

[0186] The category prediction branch is used to determine the category information in the candidate box, and its calculation formula is as follows:

[0187] .

[0188] in, To share the convolution function; For shared classification convolution kernels; It is either the Sigmoid or Softmax activation function.

[0189] The mask prediction branch is used to generate the semantic or instance segmentation mask of the target, and its calculation formula is as follows:

[0190] .

[0191] in, To share the convolution function; For shared masked convolutional kernels; The semantic or instance segmentation mask feature map output by the mask branch at scale i is used to characterize the spatial region information of the candidate target.

[0192] In this embodiment, the output , where N represents the number of mask channels, which is usually related to the number of candidate targets or the number of categories. During training, the supervision of the mask branch usually adopts binary cross-entropy (BCE) or Dice Loss.

[0193] Example 3:

[0194] The specific process in S1 is as follows:

[0195] S11: Use drones to inspect and capture images of ice accumulation on power transmission lines.

[0196] S12: Expand the dataset by performing horizontal flipping, random cropping, rotation transformation, brightness enhancement, and contrast enhancement operations on the obtained icing images of transmission lines to obtain an expanded dataset.

[0197] S13: Reintegrate the icing images of transmission lines captured in S11 with the expanded icing images of transmission lines into a new dataset as a sample library.

[0198] Example 4:

[0199] Specifically, the process of S3 is as follows:

[0200] S31: Select the icing images of the transmission lines to be labeled, and divide the icing images of the transmission lines into four labels: snow, rime, mixed rime, and ice.

[0201] In this embodiment, when annotating images of icing on transmission lines, after selecting the images of icing on transmission lines to be marked, the icing area of ​​the transmission line is accurately identified and marked, resulting in a marked box for the icing area. Based on the icing image of the transmission line in the marked box, it is divided into four labels: snow, rime, mixed rime, and ice. Finally, an area annotation box with the categories of snow, rime, mixed rime, and ice is obtained.

[0202] S32: Use the Labellmg annotation software to generate XML tag files corresponding to the icing images of transmission lines, and store the annotated sample data in Pascal VOC format.

[0203] In this embodiment, the sample data folder is VOC devkit, which includes three folders: Annotations, JPEGImages, and Image Sets.

[0204] The Annotations folder contains XML tag files for all icing on transmission lines and insulators. Each XML tag file includes an image ID, image path, image name, and the image's pixel height and width. The image's pixel height and width are represented by the four coordinates of a rectangle. , , , ,in These are the coordinates of the top-left vertex of the rectangle. These are the coordinates of the bottom right vertex of the rectangle, and each marked file corresponds to the original image in the JPEG Images.

[0205] The JPEG Images folder contains all the original, labeled images.

[0206] The Image Sets folder contains a Main subfolder. The train document in the Main subfolder contains the names of all images in the training set, and the val document contains the names of all images in the validation set.

[0207] Example 5:

[0208] Specifically, the formula for calculating the accuracy in S6 is as follows:

[0209] ,

[0210] Where P represents precision; Indicates the detected positive examples; Indicates the detected negative examples;

[0211] The formula for calculating the recall rate is:

[0212] ,

[0213] In the formula, R represents the recall rate; Indicates the detected positive examples; This indicates a positive example that was not detected.

[0214] The formula for calculating the mAP value is:

[0215] ,

[0216] in, The number of target categories involved in the calculation. Let R be the precision function under the condition of recall. Let R be the derivative with respect to recall rate R.

[0217] In this embodiment, precision reflects the proportion of cases that the model judges as positive and that are actually positive; recall reflects the proportion of test images for which the model judges as positive out of all positive test images; and mean average precision (mAP) represents the average of the average precision of each category.

[0218] Specifically, in S6, the evaluation can also be performed using frames per second, parameter quantity, and number of floating-point operations.

[0219] In this embodiment, frames per second (FPS) represents the number of image or video frames the model can process per second, and is a standard for measuring the model's processing power and efficiency. Generally, the higher the FPS, the faster the model's inference speed and the faster it processes input data. The calculation formula is as follows:

[0220] .

[0221] Where T represents the detection time for a single image; FPS is the number of images detected per second.

[0222] The number of parameters refers to the total number of parameters that a network model needs to train, affecting the model's complexity and capability. Generally, a larger number of parameters indicates a stronger fit, but it also increases the computational burden and training difficulty.

[0223] Floating point operations (FLOPs) are a metric that measures the amount of computation required for a neural network to perform one round of forward propagation. They are used to evaluate the computational complexity and performance of a model. The higher the FLOPs, the more time the model consumes.

[0224] Example 6:

[0225] In this embodiment, step S5 involves inputting the training set images partitioned in step S2 into the PCTE-Net network model obtained in step S4 for training. The specific process is as follows:

[0226] In the `train.py` file, stochastic gradient descent with a momentum of 0.9 is used for 250 training epochs. The input image pixel size is set to 640*640. The batch size is frozen for 50 epochs to 32, and then unfrozen for 200 epochs to 4. `num_workers2` is used, with the Adam optimizer, a decaying weight coefficient of 5*10⁻⁴, and an initial learning rate of 1*10⁻⁵. An IoU threshold of 0.5 is set for testing on the training set. During training, the learning rate is fine-tuned to 0.003 for better robustness. The first 50 epochs of frozen training result in rapid loss reduction, while the last 200 epochs of unfrozen training allow for continuous fine-tuning of the network. After 200 epochs, the loss on the training set gradually decreases, resulting in an optimized PCTE-Net model with the best weight data.

[0227] The training result prediction requires two files: mdal-detr.py and predict.py. First, you need to modify model_path and classes_path in mdal-detr.py. model_path points to the trained best weight file, which is located in the logs folder, and classes_path points to the txt file corresponding to the detection class. After making the modifications, you can start the prediction.

[0228] In this embodiment, Table 1 below shows the ablation experiment results of the icing detection method, and Table 2 below shows the results of different detection algorithms for the icing detection method. As can be seen from the tables, the PCTE-Net-based transmission line icing detection method achieved an mAP value of 88.1% for detecting icing, snow, rime, and mixed rime on transmission lines. This index significantly surpasses various state-of-the-art algorithms, highlighting the superior accuracy of our method. Compared with other algorithms, the model in this study has higher detection stability and anti-interference ability, and its precision and recall rate far exceed other attention mechanisms.

[0229] Table 1. Ablation test results of the icing detection method

[0230]

[0231] Table 2 Results of different detection algorithms for icing detection methods

[0232]

[0233] In this embodiment, as Figure 6As shown, the PCTE-Net-based transmission line icing detection method also performs well on heatmaps, demonstrating more accurate identification of small targets. Compared to other popular algorithms, the PCTE-Net-based method achieves higher accuracy and recall, exhibits stronger robustness, and is more suitable for icing detection of transmission lines in complex environments.

[0234] The preferred embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

[0235] Many other changes and modifications can be made without departing from the concept and scope of this invention. It should be understood that this invention is not limited to the specific embodiments, and the scope of this invention is defined by the appended claims.

Claims

1. A method for detecting icing on transmission lines based on PCTE-Net, characterized in that, Includes the following steps: S1: Acquisition of images of icing on transmission lines; S2: Divide the transmission line icing dataset obtained in S1 into a training set, a validation set, and a test set, with the ratio of the training set, validation set, and test set being 8:1:1; S3: Use data labeling tools to label the icing categories of the training set data in S2 to obtain the transmission line icing database; S4: Train the training set labeled in S3 using the improved PCTE-Net model; S5: Use the model trained in S4 to detect the data in the validation set in S2, and obtain the corresponding transmission line icing detection results; S6: The detection results in S5 are evaluated using precision, recall, and mAP value to determine the efficiency of transmission line icing detection.

2. The method for detecting icing on transmission lines based on PCTE-Net according to claim 1, characterized in that, The S4 section uses an improved PCTE-Net model based on the YOLOv11 model, with improvements including: S41: Replace the standard 3×3 convolution in the C3K2 module of the backbone with DCSConv by using point-by-point connected serpentine convolution DCSConv; S42: Add the Target Enhanced Grouped Global Normalized Attention Module (EGNA) to the part after the basic feature extraction of each scale feature is completed by C3k2; S43: Improve the original detection head using lightweight double-shared normalized convolution LWS-NormConv.

3. The method for detecting icing on transmission lines based on PCTE-Net according to claim 1, characterized in that, The specific process in S1 is as follows: S11: Using drones to inspect and capture images of ice accumulation on power transmission lines; S12: Expand the dataset by performing horizontal flipping, random cropping, rotation transformation, brightness enhancement, and contrast enhancement operations on the obtained icing images of transmission lines to obtain an expanded dataset. S13: Reintegrate the icing images of transmission lines captured in S11 with the expanded icing images of transmission lines into a new dataset as a sample library.

4. The method for detecting icing on transmission lines based on PCTE-Net according to claim 1, characterized in that, The specific process of S3 is as follows: S31: Select the icing images of the transmission lines that need to be labeled, and divide the icing images of the transmission lines into four labels: snow, rime, mixed rime, and ice; S32: Use the Labellmg annotation software to generate XML tag files corresponding to the icing images of transmission lines, and store the annotated sample data in PascalVOC format.

5. The method for detecting icing on transmission lines based on PCTE-Net according to claim 2, characterized in that, The specific process of point-by-point connected domain serpentine convolution DCSConv in S41 is as follows: S411: Let K be the coordinate set of the standard two-dimensional convolution kernel, and let its center point be... For a 3×3 convolution kernel, the coordinates can be represented as: ; S412: Apply the standard convolution kernel to... shaft and Linearize the axial directions respectively; exist On the axis, a convolution kernel of size 9, its th... Each sampling point is represented as: , , Each position is generated iteratively from the center point in sequence: , , exist Similarly, in the axial direction, it is defined as: , , in, They are respectively shaft and The offset of the axis. for The center point along the axial direction for Center point along the axis; S413: The final sampling points are obtained using bilinear interpolation. The formula for calculating the final sampling points using bilinear interpolation is as follows: , in, For the nearest integer coordinates, This is the interpolation kernel function.

6. The method for detecting icing on transmission lines based on PCTE-Net according to claim 2, characterized in that, The specific process of the Target Enhanced Grouped Global Normalized Attention Module (EGNA) in S42 is as follows: S421: Divide the input feature map into several independent groups according to channels, each group containing a portion of channels, and process them independently; S422: Add and merge all channel features within each group to extract the most critical spatial information for that group; S423: The Sigmoid function is used to compress the result to between 0 and 1, generating an attention mask; S424: Multiply the generated attention mask with the corresponding features of the initial group; S425: Re-merge all grouped enhanced features to restore the feature size to the input size, and obtain the final enhanced result.

7. The method for detecting icing on transmission lines based on PCTE-Net according to claim 2, characterized in that, The specific process of the lightweight double-shared normalized convolution LWS-NormConv in S43 is as follows: S431: Introducing a batch normalization mechanism, each input feature (P3, P4, P5) goes through the BN_Conv 1×1 module to perform linear transformation and normalization on the feature channels; S432: Introducing kernel weight sharing to reduce parameter quantity and redundancy in the core part of the detection head, the two BN_Conv 3×3 modules adopt the kernel weight sharing mechanism, that is, features of different scales use the same weight in the convolution operation; S433: Introducing a shared strategy in the detection branch, during the prediction phase, the detection head of the detection branch is divided into three parts: bounding box regression branch (Conv_Box+Scale), category classification branch (Conv_Cls), and mask prediction branch (Conv_Mask). All Conv_Box modules share a set of convolutional weights, and the Scale module learns independent scaling factors for each scale. All Conv_Cls modules share the same set of convolutional weights; All Conv_Mask modules share the same convolution kernel parameters.

8. The method for detecting icing on transmission lines based on PCTE-Net according to claim 1, characterized in that, The formula for calculating the accuracy in S6 is as follows: , Where P represents precision. For the detected positive examples, These are the detected negative examples; The formula for calculating the recall rate is: , In the formula, R is the recall rate. For the detected positive examples, These are positive examples that were not detected. The formula for calculating the mAP value is: , in, The number of target categories involved in the calculation. Let R be the precision function under the condition of recall. Let R be the derivative with respect to recall rate R.

9. The method for detecting icing on transmission lines based on PCTE-Net according to claim 1, characterized in that, The evaluation in S6 can also be based on frames per second, parameter quantity, and number of floating-point operations.

Citation Information

Cited By

  • Power transmission line icing detection method based on structural prior constraint and terminal device

    CN122368528A