Power transmission corridor overhead line fastener corrosion detection method based on deep learning
By constructing a multi-scale corrosion perception enhancement feature extraction and cross-layer adaptive pyramid feature fusion unit, the problems of low efficiency and poor accuracy in fastener corrosion detection in power transmission corridor inspection are solved, and high-precision identification and diagnosis of micro-corrosion and complex backgrounds are achieved.
Patent Information
- Application Number
- CN202511815886.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-02-27
AI Technical Summary
Existing technologies for fastener corrosion detection in power transmission corridor inspection suffer from low efficiency, poor accuracy, and insufficient robustness. In particular, they are difficult to identify minute corrosion and distinguish corrosion morphologies in complex backgrounds, and the models have weak generalization ability, making them unable to adapt to the detection needs of different scenarios.
A multi-scale corrosion perception enhancement feature extraction unit and a cross-layer adaptive pyramid feature fusion unit are constructed. Through a multi-scale attention mechanism and feature fusion module, the ability to capture and identify corrosion features is enhanced, background interference is suppressed, and high-precision positioning and diagnosis of fastener corrosion is achieved.
It significantly improves the accuracy and robustness of fastener corrosion detection, and can accurately identify minute corrosion and diverse corrosion morphologies in complex backgrounds, meeting the needs of efficient and intelligent inspection of power transmission corridors.
Smart Images

Figure CN121582676A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent operation and maintenance of power equipment, in particular to a fastener corrosion detection method for overhead lines in a power transmission corridor based on deep learning. BACKGROUND
[0002] Under the macro-background of the country's vigorous promotion of "new infrastructure" and the construction of smart power grids, the safety and reliability of the power transmission network are the lifeline of social and economic operation. Among them, the structural stability of the power transmission tower is directly related to the safe operation of the large power grid. The fasteners on the tower, such as bolts and pins, are the key nodes to ensure the structural connection and stability of the bearing capacity, and their corrosion state is one of the core indicators for evaluating the health of the tower. With the popularization of unmanned aerial vehicle inspection technology, the power transmission corridor inspection has generated a large amount of high-resolution image and video data, providing a data source for the automatic detection of a large number of fasteners, but also posing a severe challenge to the accuracy, robustness and intelligent level of the detection technology. The traditional fastener corrosion detection method with manual visual inspection and conventional image processing as the core has multiple shortcomings: High dependence on manual work and low detection efficiency: The core relies on inspection personnel to climb the tower or use high-power binoculars for observation, which is greatly affected by subjective experience and visual fatigue, and is prone to missed detection and misjudgment. Moreover, the inspection efficiency is extremely low and the labor intensity is high. Especially when crossing complex terrains such as valleys and rivers, the inspection operation risk is high, and it is difficult to meet the operation and maintenance needs of wide-area coverage and high-frequency inspection of power transmission lines.
[0003] Limited feature extraction, difficult to identify small corrosion: Conventional image processing techniques such as color space threshold segmentation, edge detection, and texture analysis require manual design of corrosion features. Such methods are extremely sensitive to environmental factors such as sudden changes in light, shadow obstruction, and surface stains, and have poor robustness. Especially for initial small rust spots or rust with similar colors to the tower body, their representation ability is severely insufficient, and the best maintenance window period is easily missed.
[0004] Single detection target, weak comprehensive diagnosis capability: Traditional methods usually treat fastener positioning and corrosion judgment as two independent and fragmented steps, and cannot achieve end-to-end detection in a unified framework. In addition, these methods cannot effectively distinguish between fastener corrosion and background interference such as red-brown tower body and bare soil, resulting in a single defect diagnosis dimension and high false positive rate.
[0005] Lack of generalization ability, poor scene adaptability: Algorithms based on manually designed features have weak generalization ability, and often need to redesign and debug parameters for different types of towers, different types of fasteners, and different seasonal lighting environments, making it difficult to reuse the operation and maintenance model and unable to achieve universal and intelligent detection of complex and variable power transmission corridor scenes.
[0006] Therefore, the power grid industry urgently needs a new intelligent detection method that integrates the strong feature self-learning capability of deep learning, the precise positioning technology of small targets, and the fine-grained state classification mechanism to achieve automated identification and accurate diagnosis of massive fasteners and their corrosion status in transmission corridor inspection images, supporting the safe, efficient, and intelligent operation and maintenance needs of the smart grid transmission link.
[0007] Currently, deep learning-based fastener corrosion detection methods are mainly divided into two categories: two-stage detection and single-stage detection. Two-stage detection methods, represented by the R-CNN series, offer advantages in detection accuracy, but their complex computational processes result in poor real-time performance, making it difficult to meet the time-sensitive requirements of power transmission corridor inspections. To improve detection efficiency, single-stage detection methods have emerged, among which the YOLO series has gained widespread attention due to its simple and efficient architecture design. Continuous iteration and optimization have significantly improved deployment efficiency while maintaining high accuracy, making it the foundational framework for current research.
[0008] To address the unique challenges of power transmission corridor inspection scenarios, researchers have made several targeted improvements and optimizations to the YOLO series models. In feature extraction, various attention mechanisms, such as channel attention (SE) and convolutional block attention (CBAM), have been introduced to enhance the model's ability to focus on critical defect areas like rust, effectively improving target recognition accuracy in complex backgrounds. For multi-scale feature processing, researchers designed architectures such as Feature Pyramid Network (FPN) and Path Aggregation Network (PAN), while also introducing an adaptive feature fusion strategy to improve the model's detection performance for fastener scale differences caused by varying shooting distances. Regarding model lightweighting, techniques such as depthwise separable convolution and lightweight backbone networks have been employed to reduce computational complexity while maintaining detection accuracy, enabling the algorithm to be deployed on edge devices such as drones.
[0009] Based on the above, existing technologies have drawbacks in the following dimensions: (1) Insufficient multi-scale capture capability of the feature extraction backbone network: During continuous downsampling, the backbone network of the existing detection model suffers severe loss of key features such as the fine-grained fastener corrosion information contained in the high-resolution shallow features, such as the outline of small fasteners in the distance and the fine texture of early rust spots, resulting in a high false negative rate for small targets. At the same time, although deep features have strong semantic information, their fixed receptive field is difficult to flexibly adapt to the drastic changes in fastener scale caused by changes in UAV flight altitude, shooting angle, etc., resulting in a decrease in the positioning accuracy of large targets in the near field.
[0010] (2) Insufficient feature fusion leads to inadequate recognition of corroded fasteners: Existing methods fail to effectively coordinate the shallow spatial details of fastener corrosion used for precise localization with the deep semantic information of fastener corrosion used for condition diagnosis when fusing multi-scale features. The coarse fusion methods such as simple addition or splicing are prone to contaminating and diluting high-resolution localization information with low-resolution category features, weakening the model's comprehensive performance in target localization and condition diagnosis, resulting in insufficient recognition of corroded fasteners.
[0011] (3) Insufficient robustness to complex background interference: The background of the transmission corridor inspection images is extremely complex, often containing a large number of towers, soil, shadows and stains that are similar in color and texture to the rust. Existing models lack effective spatial context modeling capabilities during feature extraction, which makes their attention easily interfered with and diluted by invalid background information, making it difficult to accurately focus on the fasteners and rusted areas themselves, and misjudging the background as the rusted target.
[0012] (4) Weak ability to identify intra-class differences: In reality, rust phenomena vary greatly in morphology, encompassing various forms such as point-like initial rust, sheet-like extended rust, and large-area severe rust, with different colors, textures, and shapes. Existing models lack the ability to adaptively model the diversity within the same type of defect, and their feature representation learning mechanism is relatively rigid, making it difficult to establish a unified and robust discrimination criterion for rust with different morphologies. In particular, they are prone to misclassifying atypical rust morphologies such as early yellow rust as background. Summary of the Invention
[0013] To solve the above-mentioned technical problems, the present invention provides the following technical solution: A deep learning-based method for detecting corrosion of fasteners on overhead power transmission lines includes the following specific steps: S1, Construct a dataset of images showing corrosion of fasteners in power transmission corridors; S2, input the image of corrosion of fasteners in the power transmission corridor into the multi-scale corrosion perception enhancement feature extraction unit, perform feature extraction, and output high-resolution corrosion detail feature map P1, balanced corrosion structure feature map P2, and strong semantic corrosion diagnosis feature map P3. S3. Input P1, P2, and P3 into the cross-layer adaptive pyramid feature fusion unit to perform feature fusion and obtain the detail-enhanced corrosion localization feature map P4, the balanced collaborative corrosion association feature map P5, and the semantically driven corrosion discrimination feature map P6. S4. Based on the multi-scale corrosion perception enhancement feature extraction unit and the cross-layer adaptive pyramid feature fusion unit, a multi-scale attention feature fusion corrosion detection network model MS-AFRDN (Multi-Scale AttentionFeature Fusion Rust Detection Network) is constructed for corrosion detection of overhead line fasteners in the power transmission corridor inspection scenario. The MS-AFRDN model is trained based on the power transmission corridor fastener corrosion image dataset. S5 applies the trained MS-AFRDN model to the transmission corridor inspection task to achieve intelligent identification and precise positioning of overhead line fasteners. Finally, it outputs the structured information of each detected corroded fastener, which includes the corrosion level, the bounding box for precise positioning, and the confidence score.
[0014] As a preferred embodiment of the deep learning-based method for detecting corrosion of fasteners on overhead power transmission corridors according to the present invention, the specific steps of S1 are as follows: Using drones equipped with high-definition imaging equipment, fasteners, including but not limited to bolts, nuts, and pins, at key parts of power transmission corridor towers were photographed up close. Image annotation tools were used to locate the fastener targets with rectangular bounding boxes and to label their corrosion levels, which included slight corrosion, moderate corrosion, and severe corrosion, thus constructing a data set of fastener corrosion images for power transmission corridors. Images of fastener corrosion in power transmission corridors are uniformly cropped and scaled to a fixed size. Then, data augmentation operations are performed on the images, including but not limited to random rotation, horizontal and vertical flipping, color and brightness adjustment, injection of random noise, and simulated stain occlusion, in order to expand the training sample size and improve data diversity. The dataset of images of corrosion on fasteners in power transmission corridors was divided into training, validation, and test sets according to the proportions.
[0015] As a preferred embodiment of the deep learning-based method for detecting corrosion of fasteners on overhead power transmission corridors according to the present invention, in step S2, the specific construction process of the multi-scale corrosion perception enhancement feature extraction unit is as follows: An image of fastener corrosion in an arbitrary power transmission corridor is denoted as fastener corrosion detection feature map F1. This image is used as the input to this unit. Through a CBS module consisting of a convolutional layer, a batch normalization layer, and a SiLU activation function, feature extraction and spatial downsampling are achieved, and the output is fastener corrosion detection feature map F2. F2 is then input into the CBS module to further compress the scale and extract higher-dimensional features, resulting in fastener corrosion detection feature map F3. Inputting F3 into the context-aware branch will output fastener corrosion detection feature maps F5, F8, and F11 in sequence; simultaneously, inputting F3 into the detail-focusing branch will output fastener corrosion detection feature maps F6, F9, and F12 in sequence. Perform a Concat operation on F5 and F6, and stitch them together along the channel dimension to obtain the fastener corrosion detection feature map F7; then, perform a Conv1×1 convolution operation on F7 to achieve channel integration and dimensionality reduction, and obtain a high-resolution corrosion detail feature map P1. Perform a Concat operation on F8 and F9, and stitch them together along the channel dimension to obtain the fastener corrosion detection feature map F10; then, perform a Conv1×1 convolution operation on F10 to achieve channel integration and dimensionality reduction, and obtain the balanced corrosion structure feature map P2. Perform a Concat operation on F11 and F12, concatenating them along the channel dimension to obtain the fastener corrosion detection feature map F13; then, perform a Conv1×1 convolution operation on F13 to achieve channel integration and dimensionality reduction, resulting in a strong semantic corrosion diagnosis feature map P3.
[0016] As a preferred embodiment of the deep learning-based method for detecting corrosion of fasteners on overhead power transmission corridors according to the present invention, in step S3, the specific construction process of the cross-layer adaptive pyramid feature fusion unit is as follows: P1, P2, and P3 are input into the first CHMC submodule. Through its internal cross-level modulation and calibration mechanism, multi-scale fastener corrosion information is deeply processed to obtain fastener corrosion detection feature map F14. Then, F14 is upsampled to obtain fastener corrosion detection feature map F15. Through a Concat operation, F15 and P1 are concatenated along the channel dimension to obtain fastener corrosion detection feature map F17. F17 is input into the C2F module for feature integration and transformation to obtain fastener corrosion detection feature map F18. Simultaneously, F14 is downsampled to obtain fastener corrosion detection feature map F16. Through a Concat operation, F16 and P3 are concatenated along the channel dimension to obtain fastener corrosion detection feature map F20. F20 is input into the C2F module for feature integration and transformation to obtain fastener corrosion detection feature map F21. F15 and F18 are input into the first DIFA submodule. Through a multi-branch aggregation transformation mechanism and using a lightweight depthwise separable convolution as the core operator, the two input features are efficiently interacted and fused to obtain the fastener corrosion detection feature map F19. At the same time, F16 and F21 are input into the second DIFA submodule to obtain the fastener corrosion detection feature map F22. F14, F18, and F21 are input into the second CHMC submodule for final, deepest-level cross-level aggregation and calibration, resulting in a balanced collaborative corrosion correlation feature map P5. Then, P5 is upsampled to obtain a fastener corrosion detection feature map F23. Through a Concat operation, F23 and F19 are concatenated along the channel dimension to generate a fastener corrosion detection feature map F25. F25 is input into the C2F module for final feature purification and enhancement, resulting in a detail-enhanced corrosion localization feature map P4. Simultaneously, P5 is downsampled to obtain a fastener corrosion detection feature map F24. Through a Concat operation, F24 and F22 are concatenated along the channel dimension to obtain a fastener corrosion detection feature map F26. F26 is input into the C2F module for final feature purification and enhancement, resulting in a semantically driven corrosion discrimination feature map P6.
[0017] As a preferred embodiment of the deep learning-based method for detecting corrosion of fasteners on overhead power transmission corridors according to the present invention, the specific steps of S4 are as follows: Based on the multi-scale corrosion perception enhancement feature extraction unit and the cross-layer adaptive pyramid feature fusion unit, an MS-AFRDN model is constructed for corrosion detection of overhead line fasteners in the power transmission corridor inspection scenario. The MS-AFRDN model is input into any corrosion image of a fastener in the power transmission corridor and finally outputs the structured information of the corrosion fastener, including corrosion status level, bounding box and confidence score. The MS-AFRDN model was trained based on a dataset of images of fastener corrosion in power transmission corridors.
[0018] Compared with existing technologies: When addressing the need for multi-scale corrosion detection of fasteners in power transmission corridors, this invention utilizes the synergistic effect of the context-aware branch and detail-focusing branch of the multi-scale corrosion perception enhancement feature extraction unit to accurately capture feature information of fasteners of different sizes and early micro-corrosion. This solves the problem of missed detection and inaccurate positioning of small corrosion targets on fasteners in power transmission corridors due to insufficient multi-scale target capture capability in the prior art. In response to the complex background interference in power transmission corridor inspection scenarios, this invention uses the SAFG submodule to dynamically enhance the features of key corrosion areas and suppress background noise such as tower body and shadows, effectively solving the problem of visual confusion between the target and the background and improving the salience of corrosion features. To address the issues of information conflict and dilution caused by insufficient traditional feature fusion, this invention enhances the richness and consistency of corrosion feature expression through a collaborative design of cross-level dynamic modulation calibration of the CHMC submodule and dual-stream interactive feature aggregation of the DIFA submodule. Furthermore, to address the challenges of discriminating between large intraclass differences in rust defects and the visual similarity between fastener stains and early rust, this invention enhances the model's ability to perceive subtle differences and its feature generalization through the synergistic effect of the CHMC and DIFA sub-modules. This effectively reduces the risk of misjudging similar features and ultimately significantly improves the accuracy, robustness, and reliability of fastener rust detection in complex inspection scenarios. Attached Figure Description
[0019] Figure 1 This is a flowchart of the technical solution of the present invention; Figure 2 This is a diagram of the multi-scale corrosion sensing enhancement feature extraction unit of the present invention; Figure 3 This is a structural design diagram of the SAFG submodule of the present invention; Figure 4 This is a diagram of the cross-layer adaptive pyramid feature fusion unit of the present invention; Figure 5 This is a structural design diagram of the CHMC submodule of the present invention; Figure 6 This is a structural design diagram of the DIFA submodule of the present invention; Figure 7 This is a diagram illustrating the overall architecture of the MS-AFRDN model of this invention. Figure 8 This is an image of corrosion on the fasteners of the power transmission corridor according to the present invention; Figure 9 This is an image of corrosion on fasteners in a power transmission corridor, showing the detection results of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0021] This invention provides a deep learning-based method for detecting corrosion of fasteners on overhead power transmission lines. Please refer to [link / reference]. Figures 1-9 The specific steps are as follows: S1, Construct a dataset of images showing corrosion of fasteners in power transmission corridors; The specific steps of S1 are as follows: Using drones equipped with high-definition imaging equipment, fasteners, including but not limited to bolts, nuts, and pins, at key parts of power transmission corridor towers were photographed up close. Image annotation tools were used to locate the fastener targets with rectangular bounding boxes and to label their corrosion levels, which included slight corrosion, moderate corrosion, and severe corrosion, thus constructing a data set of fastener corrosion images for power transmission corridors. Images of fastener corrosion in power transmission corridors are uniformly cropped and scaled to a fixed size. Then, data augmentation operations are performed on the images, including but not limited to random rotation, horizontal and vertical flipping, color and brightness adjustment, injection of random noise, and simulated stain occlusion, in order to expand the training sample size and improve data diversity. The dataset of images of corrosion on fasteners in power transmission corridors was divided into training, validation, and test sets according to the proportions.
[0022] Commonly used image annotation tools include LabelImg and LabelME.
[0023] Example: Automated inspections and photography along the power transmission corridor were conducted using drones equipped with high-definition cameras, acquiring a large number of on-site images of fasteners attached to critical components such as tower crossarms, node plates, and tower feet. The LabelImg professional annotation tool was used to precisely annotate the corroded fastener targets in the acquired images, using rectangular bounding boxes to accurately mark the target locations and indicate the corrosion level. The final result was a dataset of 9680 high-quality images covering three levels of defect conditions for fastener corrosion along the power transmission corridor, including 2520 images of uncorroded fasteners, 2650 images of slightly corroded fasteners, 2430 images of moderately corroded fasteners, and 2080 images of severely corroded fasteners.
[0024] Images of rusted fasteners in the power transmission corridor were uniformly cropped and scaled to 1024×1024, and then subjected to random rotation, horizontal and vertical flipping, color and brightness adjustment, injection of random noise, and simulated stain occlusion.
[0025] The dataset of images of corrosion on fasteners in power transmission corridors was divided into training, validation, and test sets in a ratio of 7:2:1.
[0026] S2, input the image of corrosion of fasteners in the power transmission corridor into the multi-scale corrosion perception enhancement feature extraction unit, perform feature extraction, and output high-resolution corrosion detail feature map P1, balanced corrosion structure feature map P2, and strong semantic corrosion diagnosis feature map P3. The main function of the multi-scale corrosion perception enhancement feature extraction unit is to adaptively capture the contextual information and local fine textures of fasteners at different scales in power transmission corridor inspection scenarios, effectively solving the problems of missed detection or inaccurate positioning caused by the diversity of target scale and corrosion features. Simultaneously, this unit can enhance the effective feature representation of fastener spatial location and corrosion areas, suppress background noise interference, and solve the problem of target and background confusion, especially in situations where fasteners and towers are visually difficult to separate, and early corrosion and surface stains are highly similar in morphology. This ensures that the method described in this invention can achieve high-precision and robust detection of key components in power transmission corridors, such as fastener corrosion. The structure of its multi-scale corrosion perception enhancement feature extraction unit is as follows: Figure 2 As shown.
[0027] In S2, the specific construction process of the multi-scale corrosion sensing enhancement feature extraction unit is as follows: An image of fastener corrosion in an arbitrary power transmission corridor is denoted as fastener corrosion detection feature map F1. This image is used as the input to this unit. Through a CBS module consisting of a convolutional layer, a batch normalization layer, and a SiLU activation function, feature extraction and spatial downsampling are achieved, and the output is fastener corrosion detection feature map F2. F2 is then input into the CBS module to further compress the scale and extract higher-dimensional features, resulting in fastener corrosion detection feature map F3. Inputting F3 into the context-aware branch outputs fastener corrosion detection feature maps F5, F8, and F11 sequentially; simultaneously, inputting F3 into the detail-focusing branch outputs fastener corrosion detection feature maps F6, F9, and F12 sequentially; the execution process of the two branches is as follows: (1) In the context-aware branch: F3 is processed by three parallel dilated convolutional layers (DilatedConv), with dilation rates r set to 1, 3 and 5 respectively, to capture context information at different scales, and output fastener corrosion detection feature maps F3-1, F3-2 and F3-3 respectively. F3-1, F3-2 and F3-3 are added element by element to fuse rich context feature information under the multi-scale receptive field to obtain fastener corrosion detection feature map F4; then, F4 is subjected to Conv3×3 convolution operation to extract features to obtain fastener corrosion detection feature map F5; then, F5 is subjected to Conv3×3 convolution operation to further extract features to obtain fastener corrosion detection feature map F8; finally, F8 is subjected to Conv3×3 convolution operation to complete the deepest level of feature extraction and output fastener corrosion detection feature map F11. (2) In the detail focus branch: F3 is input into the first SAFG submodule, and multi-scale fine-grained features are extracted through its internal multi-scale convolution kernel. Combined with the spatial attention mechanism, the features of high-value local regions are adaptively focused and enhanced to obtain the fastener corrosion detection feature map F6. Then, F6 is input into the second SAFG submodule, and the features are further enhanced through the synergistic effect of multi-scale convolution and spatial attention to output the fastener corrosion detection feature map F9. Finally, F9 is input into the third SAFG submodule to perform the final stage of feature enhancement and output the fastener corrosion detection feature map F12.
[0028] Perform a Concat operation on F5 and F6, and stitch them together along the channel dimension to obtain the fastener corrosion detection feature map F7; then, perform a Conv1×1 convolution operation on F7 to achieve channel integration and dimensionality reduction, and obtain a high-resolution corrosion detail feature map P1. Perform a Concat operation on F8 and F9, and stitch them together along the channel dimension to obtain the fastener corrosion detection feature map F10; then, perform a Conv1×1 convolution operation on F10 to achieve channel integration and dimensionality reduction, and obtain the balanced corrosion structure feature map P2. Perform a Concat operation on F11 and F12, concatenating them along the channel dimension to obtain the fastener corrosion detection feature map F13; then, perform a Conv1×1 convolution operation on F13 to achieve channel integration and dimensionality reduction, resulting in a strong semantic corrosion diagnosis feature map P3.
[0029] Example: A corrosion image of fasteners in a power transmission corridor with a size of 1024×1024×3 is input into the multi-scale corrosion perception enhancement feature extraction unit in the MS-AFRDN model, and the corrosion image of the fasteners in the power transmission corridor is denoted as the fastener corrosion detection feature map F1. F1 is input to the multi-scale corrosion sensing enhancement feature extraction unit, and through a CBS module consisting of a convolutional layer, a batch normalization layer, and a SiLU activation function, a fastener corrosion detection feature map F2 with a size of 1024×1024×64 is output; F2 is input to the CBS module, and a fastener corrosion detection feature map F3 with a size of 512×512×128 is output.
[0030] Inputting F3 into the context-aware branch outputs fastener corrosion detection feature maps F5, F8, and F11 with sizes of 256×256×128, 128×128×128, and 64×64×128 respectively. Simultaneously, inputting F3 into the detail-focusing branch outputs fastener corrosion detection feature maps F6, F9, and F12 with sizes of 256×256×256, 128×128×512, and 64×64×1024 respectively. The specific construction process of the two branches is as follows.
[0031] (1) In the context-aware branch: F3 is processed through three parallel dilated convolutional layers (DilatedConv), with dilation rates r set to 1, 3 and 5 respectively, outputting fastener corrosion detection feature maps F3-1, F3-2 and F3-3, each with a size of 512×512×128; F3-1, F3-2 and F3-3 are added element-wise, outputting a size of 512×512×128. The fastener corrosion detection feature map F4 is 256×256×128. Then, F4 is convolved with Conv3×3 to output a fastener corrosion detection feature map F5 with a size of 256×256×128. Next, F5 is convolved with Conv3×3 to output a fastener corrosion detection feature map F8 with a size of 128×128×128. Finally, F8 is convolved with Conv3×3 to output a fastener corrosion detection feature map F11 with a size of 64×64×128. (2) In the detail focus branch: F3 is input into the first SAFG submodule for processing, and the output fastener corrosion detection feature map F6 with a size of 256×256×256 is output; then, F6 is input into the second SAFG submodule for processing, and the output fastener corrosion detection feature map F9 with a size of 128×128×512 is output; finally, F9 is input into the third SAFG submodule for processing, and the output fastener corrosion detection feature map F12 with a size of 64×64×1024 is output. Perform a Concat operation on F5 and F6, stitching them along the channel dimension to output a fastener corrosion detection feature map F7 with a size of 256×256×384; perform a Conv1×1 convolution operation on F7 to output a high-resolution corrosion detail feature map P1 with a size of 256×256×256. Perform a Concat operation on F8 and F9, stitching them along the channel dimension to output a fastener corrosion detection feature map F10 with a size of 128×128×640; perform a Conv1×1 convolution operation on F10 to output a balanced corrosion structure feature map P2 with a size of 128×128×512. Perform a Concat operation on F11 and F12, concatenating them along the channel dimension to output a fastener corrosion detection feature map F13 with a size of 64×64×1152; perform a Conv1×1 convolution operation on F13 to output a strong semantic corrosion diagnosis feature map P3 with a size of 64×64×1024.
[0032] This invention designs a fine-grained spatial enhancement module (SAFG). The SAFG submodule employs a feature enhancement design that combines multi-scale detail perception with spatial attention. Its core function is to adaptively focus on and enhance local regions relevant to the detection task, addressing issues such as small fastener target scale, rich rust details, and complex background interference in power transmission corridor inspection images, thereby improving the ability to accurately identify rusted fasteners.
[0033] The SAFG submodule employs a parallel dual-branch design: the first branch is for multi-scale fine-grained feature extraction, which uses multi-scale convolution to capture detailed texture information under different receptive fields; the second branch is for spatial attention weight generation, which aggregates global spatial information of corrosion features through global average pooling (GAP) and global max pooling (GMP) along the common dimension, and then generates a spatial attention weight map through convolution and a sigmoid activation function. Finally, by using an element-wise multiplication strategy between the multi-scale fine-grained features and the spatial attention weight map, dynamic weighting and enhancement of key corrosion feature regions are achieved while suppressing background interference. The SAFG submodule effectively integrates multi-scale detailed information and spatial attention mechanisms, enhancing adaptive focusing and feature enhancement of salient regions of corrosion detection targets, thereby improving the feature extraction effect for fasteners of various scales, especially small targets and early micro-corrosion. The structure diagram of the SAFG submodule is shown below. Figure 3 As shown.
[0034] The construction and execution process of the SAFG submodule is as follows: A Conv3×3 convolution operation is performed on the fine-grained spatial focus enhancement feature map X1 to achieve preliminary feature extraction and downsampling, and the output is the fine-grained spatial focus enhancement feature map X2; a Conv1×1 convolution operation is performed on X2 to achieve channel adjustment and information integration, and the output is the fine-grained spatial focus enhancement feature map X3. Subsequently, X3 is input into the multi-scale fine-grained feature extraction branch, outputting a fine-grained spatial focusing enhancement feature map X7; simultaneously, X3 is input into the spatial attention weight generation branch, outputting a spatial attention weight map W1. The specific construction process of the two branches is as follows: (1) In the multi-scale fine-grained feature extraction branch, Conv3×3, Conv5×5 and Conv7×7 convolution operations are performed on X3 in parallel to capture feature information under different receptive fields, and fine-grained spatial focusing enhancement feature map X4, fine-grained spatial focusing enhancement feature map X5 and fine-grained spatial focusing enhancement feature map X6 are obtained respectively; then, X4, X5 and X6 are added element by element to realize the fusion of multi-scale fastener corrosion information and obtain fine-grained spatial focusing enhancement feature map X7. (2) In the spatial attention weight generation branch, X3 is input in parallel into the global average pooling (GAP) and global max pooling (GMP) layers, and pooling is performed along the channel dimension to aggregate the global context information of the rust feature, resulting in fine-grained spatial focus enhancement feature maps X8 and X9, respectively. X8 and X9 are then concatenated along the channel dimension to obtain fine-grained spatial focus enhancement feature map X10. Subsequently, X10 is subjected to a Conv7×7 convolution operation to effectively aggregate local spatial context information and compress the channels, resulting in a single-channel fine-grained spatial focus enhancement feature map X11. Finally, X11 is processed using the Sigmoid activation function to obtain a weight vector map with the importance of each channel, i.e., the spatial attention weight map W1. Finally, X7 and W1 are multiplied element-wise to achieve dynamic weighting and enhancement of key rust feature regions, highlighting high-value rust information and suppressing irrelevant background, resulting in a fine-grained spatial focusing enhanced feature map X12. X12 is then input into the C2F module for deep rust feature fusion and enhancement of expressive power, resulting in a fine-grained spatial focusing enhanced feature map X13.
[0035] Example: The construction and execution process of the SAFG submodule is as follows, taking the first SAFG submodule as an example: input F3 into the first SAFG submodule, and denote F3 as the fine-grained spatial focusing enhancement feature map X1.
[0036] First, perform a Conv3×3 convolution operation on X1 to output a fine-grained spatial focus enhancement feature map X2 with a size of 256×256×128; then perform a Conv1×1 convolution operation on X2 to output a fine-grained spatial focus enhancement feature map X3 with a size of 256×256×256.
[0037] Subsequently, X3 is input into the multi-scale fine-grained feature extraction branch, outputting a fine-grained spatial focusing enhancement feature map X7 with a size of 256×256×256; simultaneously, X3 is input into the spatial attention weight generation branch, outputting a spatial attention weight map W1 with a size of 256×256×1. The specific construction process of the two branches is as follows: (1) In the multi-scale fine-grained feature extraction branch, Conv3×3, Conv5×5 and Conv7×7 convolution operations are performed on X3 in parallel, and the output sizes are all 256×256×256 fine-grained spatial focus enhancement feature maps X4, X5 and X6; then, X4, X5 and X6 are added element by element, and the output size is 256×256×256 fine-grained spatial focus enhancement feature map X7.
[0038] (2) In the spatial attention weight generation branch, X3 is input into the Global Average Pooling (GAP) and Global Max Pooling (GMP) layers in parallel and pooled along the channel dimension to output fine-grained spatial focus enhancement feature maps X8 and X9, both with a size of 256×256×1. X8 and X9 are concatenated along the channel dimension to output fine-grained spatial focus enhancement feature map X10 with a size of 256×256×2. Subsequently, X10 is convolved with Conv7×7 to output fine-grained spatial focus enhancement feature map X11 with a size of 256×256×1. Finally, X11 is processed by the Sigmoid activation function to output spatial attention weight map W1 with a size of 256×256×1.
[0039] Finally, X7 and W1 are multiplied element-wise to output a fine-grained spatial focus enhancement feature map X12 with a size of 256×256×256. X12 is then input into the C2F module for processing to output a fine-grained spatial focus enhancement feature map X13 with a size of 256×256×256.
[0040] The CBS module consists of convolutional layers, batch normalization, and the SiLU activation function. This module extracts features by performing convolutional operations on the input feature map, uses a stride of 2 for downsampling to compress the feature map size and expand the receptive field, accelerates the training process and improves stability, and the SiLU activation function further enhances non-linear expressive power.
[0041] Dilated convolution is a convolution operation that expands the receptive field by introducing gaps between elements of the convolution kernel, enabling the capture of a wider range of contextual information without increasing computational cost or the number of parameters.
[0042] The Conv3×3 convolution operation represents the process of processing a feature image using a convolution kernel of size 3×3; the Conv1×1 convolution operation represents the process of processing a feature image using a convolution kernel of size 1×1.
[0043] The C2F module is a highly efficient cross-stage feature fusion module, whose core consists of n bettleneck structures. This module divides the input features into two parts through a split operation. One part undergoes rich feature transformation and extraction through multiple bettlenecks, while the other part retains the original feature information. Finally, the transformed features are fused with the original features through concatenation and convolution operations.
[0044] S3. Input P1, P2, and P3 into the cross-layer adaptive pyramid feature fusion unit to perform feature fusion and obtain the detail-enhanced corrosion localization feature map P4, the balanced collaborative corrosion association feature map P5, and the semantically driven corrosion discrimination feature map P6. The cross-layer adaptive pyramid feature fusion unit, through the scheduling and fusion of multi-level features, simultaneously achieves precise positioning of key components and corrosion level diagnosis information in power transmission corridor inspection scenarios. This unit can generate more accurate and highly recognizable fastener corrosion fusion feature representations, effectively avoiding dilution or conflict between corrosion positioning and diagnostic information during fusion. Facing diverse morphologies and complex backgrounds of corrosion targets, it can suppress false feature interference, maintain high recognizability for highly similar features such as early point corrosion and stains, and robustly identify significant intra-class differences in different corrosion levels, achieving high-precision and robust detection of fastener corrosion in power transmission corridors. The structure of the cross-layer adaptive pyramid feature fusion unit is as follows: Figure 4 As shown.
[0045] In S3, the specific construction process of the cross-layer adaptive pyramid feature fusion unit is as follows: P1, P2, and P3 are input into the first CHMC submodule. Through its internal cross-level modulation and calibration mechanism, multi-scale fastener corrosion information is deeply processed to obtain fastener corrosion detection feature map F14. Then, F14 is upsampled to obtain fastener corrosion detection feature map F15. Through a Concat operation, F15 and P1 are concatenated along the channel dimension to obtain fastener corrosion detection feature map F17. F17 is input into the C2F module for feature integration and transformation to obtain fastener corrosion detection feature map F18. Simultaneously, F14 is downsampled to obtain fastener corrosion detection feature map F16. Through a Concat operation, F16 and P3 are concatenated along the channel dimension to obtain fastener corrosion detection feature map F20. F20 is input into the C2F module for feature integration and transformation to obtain fastener corrosion detection feature map F21. F15 and F18 are input into the first DIFA submodule. Through a multi-branch aggregation transformation mechanism and using a lightweight depthwise separable convolution as the core operator, the two input features are efficiently interacted and fused to obtain the fastener corrosion detection feature map F19. At the same time, F16 and F21 are input into the second DIFA submodule to obtain the fastener corrosion detection feature map F22. F14, F18, and F21 are input into the second CHMC submodule for final, deepest-level cross-level aggregation and calibration, resulting in a balanced collaborative corrosion correlation feature map P5. Then, P5 is upsampled to obtain a fastener corrosion detection feature map F23. Through a Concat operation, F23 and F19 are concatenated along the channel dimension to generate a fastener corrosion detection feature map F25. F25 is input into the C2F module for final feature purification and enhancement, resulting in a detail-enhanced corrosion localization feature map P4. Simultaneously, P5 is downsampled to obtain a fastener corrosion detection feature map F24. Through a Concat operation, F24 and F22 are concatenated along the channel dimension to obtain a fastener corrosion detection feature map F26. F26 is input into the C2F module for final feature purification and enhancement, resulting in a semantically driven corrosion discrimination feature map P6.
[0046] Example: P1, P2, and P3 are input into the cross-layer adaptive pyramid feature fusion unit for feature fusion, and the following outputs are respectively: detail-enhanced rust localization feature map P4, balanced collaborative rust association feature map P5, and semantically driven rust discrimination feature map P6 with sizes of 256×256×256, 128×128×512, and 64×64×1024.
[0047] P1, P2, and P3 are input together into the first CHMC submodule for processing, outputting a fastener corrosion detection feature map F14 with a size of 128×128×512. Then, F14 is upsampled, outputting a fastener corrosion detection feature map F15 with a size of 256×256×512. Through the Concat operation, F15 and P1 are concatenated along the channel dimension, outputting a fastener corrosion detection feature map F17 with a size of 256×256×768. F17 is input into the C2F module for processing, outputting a fastener corrosion detection feature map F18 with a size of 256×256×768. Meanwhile, F14 is downsampled to output a fastener corrosion detection feature map F16 with a size of 64×64×512; F16 and P3 are concatenated along the channel dimension through a Concat operation to output a fastener corrosion detection feature map F20 with a size of 64×64×1536; F20 is input into the C2F module for processing to output a fastener corrosion detection feature map F21 with a size of 64×64×1536. F15 and F18 are input together into the first DIFA submodule for processing, and the output fastener corrosion detection feature map F19 with a size of 256×256×256 is output; at the same time, F16 and F21 are input together into the second DIFA submodule for processing, and the output fastener corrosion detection feature map F22 with a size of 64×64×1024 is output. F14, F18, and F21 are input together into the second CHMC submodule for processing, outputting a balanced collaborative corrosion correlation feature map P5 with a size of 128×128×512. Then, P5 is upsampled, outputting a fastener corrosion detection feature map F23 with a size of 256×256×512. Through a Concat operation, F23 and F19 are concatenated along the channel dimension, outputting a fastener corrosion detection feature map F25 with a size of 256×256×768. F25 is then input into the C2F module for processing. The system outputs a detailed enhanced rust localization feature map P4 with a size of 256×256×256. Simultaneously, P5 is downsampled to output a fastener rust detection feature map F24 with a size of 64×64×512. Through a Concat operation, F25 and F22 are concatenated along the channel dimension to output a fastener rust detection feature map F26 with a size of 64×64×1536. F26 is then input into the C2F module for processing to output a semantically driven rust discrimination feature map P6 with a size of 64×64×1024.
[0048] This invention designs a Cross-Layer Feature Modulation and Calibration Module (CHMC). The CHMC submodule is a feature processing structure designed to achieve multi-scale feature depth collaboration. Its core function is to realize cross-layer dynamic feature modulation to address the insufficient recognition capability of rusted fasteners caused by inadequate traditional feature fusion. The core mechanism of this submodule is as follows: by receiving rust feature inputs from high, medium, and low layers in parallel, processing the high-level semantic features and low-level detail features separately, and generating two dynamic modulation weights using the Sigmoid function; subsequently, using these two weights to perform element-wise multiplication on the main feature stream of the intermediate layer, thereby enabling deep semantic information to guide shallow spatial details, while simultaneously using shallow details to calibrate the representation of deep features. Through this dynamic modulation mechanism, the CHMC submodule can generate more abstract and generalizable fused rust features, improving the model's ability to generalize and recognize rust defects of various morphologies. The specific structural design of the CHMC submodule is as follows. Figure 5 Show.
[0049] The three feature maps input to the CHMC submodule are denoted as cross-level modulation and calibration feature map Y1, cross-level modulation and calibration feature map Y2, and cross-level modulation and calibration feature map Y3, respectively. The construction and execution process of the CHMC submodule is as follows [the steps of the CHMC submodule are denoted as A1, A2, A3...]: A1. Perform downsampling on Y1 to obtain cross-level modulation and calibration feature map Y4; then, perform Conv3×3 convolution on Y4 to extract key semantic information of fastener corrosion, and obtain cross-level modulation and calibration feature map Y5; finally, generate high-level semantic modulation weights O1 through the Sigmoid activation function. A2. Input Y2 in parallel into three convolutional branches for deep feature extraction and enhancement: (1) In the first branch, perform Conv3×3 convolution operation on Y2 to obtain cross-level modulation and calibration feature map Y6; (2) In the second branch, two consecutive Conv3×3 convolution operations are performed on Y2 to enhance the ability to express complex patterns such as rust by adding nonlinear transformation, and to construct a hierarchical and more efficient feature representation with fewer computational parameters, thereby obtaining cross-level modulation and calibration feature map Y7 and cross-level modulation and calibration feature map Y8 in sequence. (3) In the third branch, Y2 and Y8 are multiplied element by element to achieve self-calibration and enhancement of features, and cross-level modulation and calibration feature map Y9 is obtained; then, Y9 and Y6 are added element by element to fuse the original information with the information after deep extraction, and cross-level modulation and calibration feature map Y10 is obtained.
[0050] Subsequently, a Conv1×1 convolution operation was performed on Y10 to interact and integrate channel information, resulting in a cross-level modulation and calibration feature map Y11. A3. Upsample Y3 to obtain cross-level modulation and calibration feature map Y12; then perform Conv3×3 convolution on Y12 to extract key details of fastener corrosion and obtain cross-level modulation and calibration feature map Y13; finally, generate low-level detail modulation weights O2 through the Sigmoid activation function. A4. Multiply O1 and Y11 element-wise to enhance and calibrate Y11 using high-level semantic information, resulting in a cross-level modulation and calibration feature map Y14. Simultaneously, multiply O2 and Y11 element-wise to supplement and calibrate Y11 using low-level detail information, resulting in a cross-level modulation and calibration feature map Y15. Subsequently, perform a Concat operation on Y11, Y14, and Y15, concatenating them along the channel dimension to fully aggregate the feature information enhanced in different ways from the three paths, resulting in a cross-level modulation and calibration feature map Y16. A5. Input Y16 into the CBS module for final feature purification and integration to obtain the cross-level modulation and calibration feature map Y17.
[0051] Example (taking the first CHMC submodule as an example): Input P1, P2, and P3 into the first CHMC submodule, and denote P1, P2, and P3 as cross-level modulation and calibration feature map Y1, cross-level modulation and calibration feature map Y2, and cross-level modulation and calibration feature map Y3; A1. Downsample Y1 to output a cross-level modulation and calibration feature map Y4 with a size of 128×128×256; then, perform a Conv3×3 convolution operation on Y4 to output a cross-level modulation and calibration feature map Y5 with a size of 128×128×512; finally, pass the Sigmoid activation function to output a high-level semantic modulation weight O1 with a size of 128×128×512. A2. Input Y2 in parallel into three convolutional branches for deep feature extraction and enhancement: (1) In the first branch, perform Conv3×3 convolution operation on Y2 to output a cross-level modulation and calibration feature map Y6 with a size of 128×128×512; (2) In the second branch, two consecutive Conv3×3 convolution operations are performed on Y2 to output cross-level modulation and calibration feature maps Y7 and Y8, both with a size of 128×128×512. (3) In the third branch, Y2 and Y8 are multiplied element by element to output a cross-level modulation and calibration feature map Y9 with a size of 128×128×512; then, Y9 and Y6 are added element by element to output a cross-level modulation and calibration feature map Y10 with a size of 128×128×512.
[0052] Subsequently, a Conv1×1 convolution operation is performed on Y10 to output a cross-level modulation and calibration feature map Y11 with a size of 128×128×512; A3. Upsample Y3 to output a cross-level modulation and calibration feature map Y12 with a size of 128×128×1024; then perform a Conv3×3 convolution operation on Y12 to output a cross-level modulation and calibration feature map Y13 with a size of 128×128×512; finally, generate a low-level detail modulation weight O2 with a size of 128×128×512 through the Sigmoid activation function. A4. Multiply O1 and Y11 element-wise to output a cross-level modulation and calibration feature map Y14 with a size of 128×128×512; simultaneously, multiply O2 and Y11 element-wise to output a cross-level modulation and calibration feature map Y15 with a size of 128×128×512; then, perform a Concat operation on Y11, Y14, and Y15, concatenating them along the channel dimension to output a cross-level modulation and calibration feature map Y16 with a size of 128×128×512. A5. Input Y16 into the CBS module for processing, and output a cross-level modulation and calibration feature map Y17 with a size of 128×128×512.
[0053] This invention designs a Cross-Stream Gating and Fusion Module (DIFA). The DIFA submodule is a feature processing structure designed to achieve deep interaction and enhancement of multi-channel features, aiming to address the problem of insufficient feature representation capabilities in power transmission corridor inspection scenarios, leading to inadequate discrimination of visually similar states such as surface stains and early corrosion on fasteners. The core mechanism of this submodule lies in a two-order cross-residual fusion paradigm: after preliminary feature transformation of the feature maps received in parallel from two input streams, information interaction across the input streams is immediately performed as residual supplementation; simultaneously, a multi-channel parallel architecture employing decoupled spatial and channel transformations is used to deeply refine the features of each input stream. Through this two-order deep interaction method, more abstract and generalizable fused corrosion features can be generated, enhancing the model's ability to perceive fine details and subtle differences, and strengthening the generalized recognition ability of corrosion defects with diverse morphologies, thereby effectively reducing misjudgments caused by blurred visual features. The specific structural design of the DIFA submodule is as follows... Figure 6 As shown.
[0054] The input feature maps of the two input streams of the DIFA submodule are considered as two-stream interactive aggregated feature maps Z1 and Z2. The construction and execution process of the DIFA submodule is as follows: [The steps of the DIFA submodule are denoted as B1, B2, B3...]: B1. Input Z1 into the upper input stream of the DIFA submodule for processing, outputting a dual-stream interactive aggregated feature map Z17; input Z2 into the lower input stream of the DIFA submodule for processing, outputting a dual-stream interactive aggregated feature map Z18. The specific construction process of the two input streams is as follows: (1) In the upper input stream, Z1 is subjected to Conv1×1 convolution operation to adjust the number of channels and perform feature transformation to obtain the dual-stream interactive aggregated feature map Z3; Z3 is subjected to Split operation to divide it into two parts along the channel dimension to obtain the dual-stream interactive aggregated feature map Z4 and the dual-stream interactive aggregated feature map Z6 respectively; Z4 is subjected to DWConv3×3 convolution operation to extract features to obtain the dual-stream interactive aggregated feature map Z5; at the same time, Z6 is subjected to two consecutive DWConv3×3 convolution operations to obtain the dual-stream interactive aggregated feature map Z7 and the dual-stream interactive aggregated feature map Z8 respectively; then, Z5 and Z8 are subjected to Concat operation to concatenate along the channel dimension to obtain the dual-stream interactive aggregated feature map Z9; finally, Z9 is added element-wise to Z10 from the lower input stream to achieve cross-input stream information supplementation and enhancement to obtain the dual-stream interactive aggregated feature map Z17.
[0055] (2) In the lower input stream, perform a Conv1×1 convolution operation on Z2 to adjust the number of channels and perform feature transformation to obtain the dual-stream interactive aggregated feature map Z10; perform a Split operation on Z10 to divide it into two parts along the channel dimension to obtain the dual-stream interactive aggregated feature map Z11 and the dual-stream interactive aggregated feature map Z13 respectively; perform a DWConv3×3 convolution operation on Z11 to extract features to obtain the dual-stream interactive aggregated feature map Z12; at the same time, perform two consecutive DWConv3×3 convolution operations on Z13 to obtain the dual-stream interactive aggregated feature map Z14 and the dual-stream interactive aggregated feature map Z15 in sequence; then, perform a Concat operation on Z12 and Z15 to concatenate them along the channel dimension to obtain the dual-stream interactive aggregated feature map Z16; add Z16 to Z3 from the upper input stream element by element to achieve information supplementation and enhancement across the input stream to obtain the dual-stream interactive aggregated feature map Z18.
[0056] B2. Perform a Concat operation on the refined and enhanced feature maps Z17 and Z18 from the two input streams, concatenating them along the channel dimension to fully integrate all the corrosion feature information from the two input streams, resulting in a dual-stream interactive aggregated feature map Z19. Then, perform a Conv1×1 convolution operation on Z19 for final feature purification and channel integration, resulting in a dual-stream interactive aggregated feature map Z20.
[0057] Example (taking the first DIFA submodule as an example): Input F15 and F18 into the first DIFA submodule, and denote F15 and F18 as the two-stream interactive aggregation feature map Z1 and the two-stream interactive aggregation feature map Z2; B1. Input Z1 into the upper input stream for processing, outputting a dual-stream interactive aggregated feature map Z17 with a size of 256×256×256; input Z2 into the lower input stream for processing, outputting a dual-stream interactive aggregated feature map Z18 with a size of 256×256×256. The specific construction process of the two input streams is as follows: (1) In the upper input stream, perform a Conv1×1 convolution operation on Z1 to output a dual-stream interactive aggregated feature map Z3 with a size of 256×256×256; perform a Split operation on Z3 to divide it into two parts along the channel dimension, and output dual-stream interactive aggregated feature maps Z4 and Z6 with a size of 256×256×128 respectively; perform a DWConv3×3 convolution operation on Z4 to output a dual-stream interactive aggregated feature map Z5 with a size of 256×256×128. Simultaneously, Z6 is subjected to two consecutive DWConv3×3 convolution operations, outputting two dual-stream interactive aggregated feature maps Z7 and Z8, each with a size of 256×256×128. Subsequently, Z5 and Z8 are concatted, outputting a dual-stream interactive aggregated feature map Z9 with a size of 256×256×256. Finally, Z9 is added element-wise to Z10 from the lower input stream, outputting a dual-stream interactive aggregated feature map Z17 with a size of 256×256×256.
[0058] (2) In the lower input stream, perform a Conv1×1 convolution operation on Z2 to output a dual-stream interactive aggregated feature map Z10 with a size of 256×256×256; perform a Split operation on Z10 to divide it into two parts along the channel dimension, outputting dual-stream interactive aggregated feature maps Z11 and Z13 with a size of 256×256×128 respectively; perform a DWConv3×3 convolution operation on Z11 to output a dual-stream interactive aggregated feature map Z1 with a size of 256×256×128. 2; Simultaneously, Z13 is stacked with two consecutive DWConv3×3 convolution operations, outputting two dual-stream interactive aggregated feature maps Z14 and Z15, each with a size of 256×256×128; then, Z12 and Z15 are concatted, outputting a dual-stream interactive aggregated feature map Z16 with a size of 256×256×256; finally, Z16 is added element-wise to Z3 from the upper input stream, outputting a dual-stream interactive aggregated feature map Z18 with a size of 256×256×256.
[0059] B2. Perform a Concat operation on Z17 and Z18, concatenating them along the channel dimension to output a two-stream interactive aggregated feature map Z19 with a size of 256×256×512. Then perform a Conv1×1 convolution operation on Z19 to output a two-stream interactive aggregated feature map Z20 with a size of 256×256×256.
[0060] in: Depthwise separable convolution (DWConv) is a lightweight convolution operation that breaks down standard convolution into two steps: depthwise convolution and pointwise convolution, which can significantly reduce computation and the number of parameters.
[0061] S4, based on a multi-scale corrosion perception enhancement feature extraction unit and a cross-layer adaptive pyramid feature fusion unit, constructs a multi-scale attention feature fusion corrosion detection network model MS-AFRDN (Multi-Scale AttentionFeature Fusion Rust Detection Network) for corrosion detection of overhead line fasteners in power transmission corridor inspection scenarios. The MS-AFRDN model is trained using a dataset of power transmission corridor fastener corrosion images. The overall architecture of the MS-AFRDN model is as follows: Figure 7 As shown.
[0062] The specific steps of S4 are as follows: Based on a multi-scale corrosion perception enhancement feature extraction unit and a cross-layer adaptive pyramid feature fusion unit, an MS-AFRDN model is constructed for corrosion detection of overhead line fasteners in power transmission corridor inspection scenarios. An image of any corrosion fastener in a power transmission corridor is input into the MS-AFRDN model, and the model outputs the structured information of the corroded fastener, including the corrosion state level, bounding box, and confidence score; as shown below: The image of corrosion of fasteners in any power transmission corridor is input into the multi-scale corrosion perception enhancement feature extraction unit to extract features and output a high-resolution corrosion detail feature map P1, a balanced corrosion structure feature map P2, and a strong semantic corrosion diagnosis feature map P3. P1, P2, and P3 are input into the cross-layer adaptive pyramid feature fusion unit to perform feature fusion and output the detail-enhanced corrosion localization feature map P4, the balanced collaborative corrosion association feature map P5, and the semantically driven corrosion discrimination feature map P6. Inputting P4, P5, and P6 into the detection head yields diagnostic information results including fastener corrosion level, bounding box coordinates, and confidence score. After feature fusion of these results, valid prediction boxes are selected based on the confidence threshold, and then duplicate boxes are eliminated using the Soft Non-Maximum Suppression (Soft-NMS) algorithm. The resulting image is then labeled with corrosion level, bounding box, and corresponding confidence score.
[0063] Example: Let the dimensions of each of the following images or feature maps be H×W×C, where H is the height of the feature map, W is the width of the feature map, and C is the number of channels of the feature map.
[0064] First, the corrosion image of fasteners in the power transmission corridor with a size of 1024×1024×3 is denoted as the fastener corrosion detection feature map F1. It is input into the multi-scale corrosion perception enhancement feature extraction unit in the MS-AFRDN model for feature extraction. The output is a high-resolution corrosion detail feature map P1 with a size of 256×256×256, a balanced corrosion structure feature map P2 with a size of 128×128×512, and a strong semantic corrosion diagnosis feature map P3 with a size of 64×64×1024. Subsequently, P1, P2, and P3 are input into the cross-layer adaptive pyramid feature fusion unit for feature fusion, and output a detail-enhanced rust localization feature map P4 with a size of 256×256×256, a balanced collaborative rust association feature map P5 with a size of 128×128×512, and a semantically driven rust discrimination feature map P6 with a size of 64×64×1024. Finally, P4, P5, and P6 are input into the subsequent detection head to obtain diagnostic information results of fastener corrosion level, bounding box coordinates, and confidence score; feature fusion is performed on the results, and soft nonmaximum suppression (Soft-NMS) is used for filtering to output an image labeled with corrosion level, bounding box, and corresponding confidence score.
[0065] The detection head, located in the final stage of the model, primarily transforms the deep features extracted by the backbone network and feature fusion module into specific detection results. It typically includes a classification branch to determine the target category and a regression branch to predict the target's location and size in the image. In some tasks, it can also be extended to output target confidence, key points, or segmentation masks. The head is the core part of the network that directly generates detection or segmentation results.
[0066] Soft Nonmaximum Suppression (Soft-NMS) is an optimization algorithm in the post-processing stage of object detection, primarily designed to address the issue of missed detections of dense targets caused by the hard suppression of traditional NMS. It calculates the Intersection over Union (IoU) of candidate boxes and employs a confidence-weighted attenuation strategy for highly overlapping boxes, with more significant attenuation as the overlap increases. In multi-object scenarios, it iteratively preserves the optimal bounding boxes for each object, avoiding redundant annotations and improving detection robustness. Soft-NMS is a core post-processing step for optimizing the accuracy of dense object detection.
[0067] The MS-AFRDN model was trained based on a dataset of images showing corrosion of fasteners in power transmission corridors. The specific model training process is as follows: Before training the MS-AFRDN model, all weight parameters of the MS-AFRDN model are initialized, and corresponding hyperparameters, including learning rate, batch size, optimizer type, and momentum, are configured according to the characteristics of the fastener corrosion detection task. After setting the hyperparameters, a batch iterative training strategy is adopted. Each training batch inputs a certain number of labeled fastener corrosion images from the transmission corridor into the MS-AFRDN model, which then sequentially passes through a multi-scale corrosion perception enhancement feature extraction unit, a cross-layer adaptive pyramid feature fusion unit, and a detection head for forward propagation to obtain the predicted fastener location information and corrosion target level of the transmission corridor. Subsequently, the error between the predicted results and the actual labels is calculated using a loss function, and the network weights are dynamically updated using a backpropagation algorithm and gradient descent strategy to optimize model performance. During training, an independent validation set is used to monitor model performance in real time. Changes in the mean accuracy and loss value on the validation set are used to determine if the model is overfitting. If the validation set performance does not improve within a certain period, the learning rate is dynamically adjusted or an early stopping mechanism is enabled to prevent overfitting. Additionally, a test set is used for the final evaluation of model performance. When the validation set loss no longer decreases significantly after several consecutive training rounds or the model performance stabilizes, the training process is terminated; finally, the model weights and parameters that perform best on the validation set are saved to obtain the trained MS-AFRDN model that can be used for high-precision corrosion detection of overhead line fasteners.
[0068] S5 applies the trained MS-AFRDN model to the transmission corridor inspection task to achieve intelligent identification and precise positioning of overhead line fasteners. Finally, it outputs the structured information of each detected corroded fastener, which includes the corrosion level, the bounding box for precise positioning, and the confidence score.
[0069] Example: By using drones equipped with high-definition imaging devices to conduct automated inspections and photography along the power transmission corridor, the system acquires a large number of on-site images of fasteners attached to key parts such as tower crossarms and tower feet. Figure 8 As shown, the image of fastener corrosion in the power transmission corridor is input into a multi-scale attention feature fusion corrosion detection network model, and the output image includes the fastener corrosion level, bounding box, and confidence score, as shown. Figure 9 As shown.
[0070] Figure 9 The test results for the four parts are as follows: (1) Corrosion level: slight corrosion, confidence level 0.89, coordinates (102,665,307,870) (2) Corrosion level: slight corrosion, confidence level 0.91, coordinates (665,358,870,563) (3) Corrosion level: moderate corrosion, confidence level 0.92, coordinates (563,696,768,901) (3) Corrosion level: severe corrosion, confidence level 0.91, coordinates (410,410,614,614).
[0071] Although the present invention has been described above with reference to embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of the invention. In particular, as long as there is no structural conflict, the features in the disclosed embodiments can be combined with each other in any manner. The lack of an exhaustive description of these combinations in this specification is merely for the sake of brevity and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A method for detecting corrosion of fasteners on overhead power transmission lines based on deep learning, characterized in that, The specific steps are as follows: S1, Construct a dataset of images showing corrosion of fasteners in power transmission corridors; S2, input the image of corrosion of fasteners in the power transmission corridor into the multi-scale corrosion perception enhancement feature extraction unit, perform feature extraction, and output high-resolution corrosion detail feature map P1, balanced corrosion structure feature map P2, and strong semantic corrosion diagnosis feature map P3. S3. Input P1, P2, and P3 into the cross-layer adaptive pyramid feature fusion unit to perform feature fusion and obtain the detail-enhanced corrosion localization feature map P4, the balanced collaborative corrosion association feature map P5, and the semantically driven corrosion discrimination feature map P6. S4. Based on the multi-scale corrosion perception enhancement feature extraction unit and the cross-layer adaptive pyramid feature fusion unit, a multi-scale attention feature fusion corrosion detection network model MS-AFRDN is constructed for corrosion detection of overhead line fasteners in the power transmission corridor inspection scenario. The MS-AFRDN model is trained based on the power transmission corridor fastener corrosion image dataset. S5 applies the trained MS-AFRDN model to the power transmission corridor inspection task, realizing intelligent identification and precise positioning of overhead line fasteners, and finally outputting the structured information of each detected corroded fastener.
2. The method for detecting corrosion of fasteners on overhead power transmission lines based on deep learning according to claim 1, characterized in that, The specific steps of S1 are as follows: A drone equipped with high-definition imaging equipment was used to take close-up photos of fasteners on key parts of the transmission corridor towers; an image annotation tool was used to locate the fastener targets with rectangular bounding boxes and to annotate their corrosion levels, thus constructing a dataset of fastener corrosion images for the transmission corridor. Images of fastener corrosion in power transmission corridors are uniformly cropped and scaled to a fixed size, and then data augmentation operations are performed on the images to expand the training sample size and improve data diversity. The dataset of images of corrosion on fasteners in power transmission corridors was divided into training, validation, and test sets according to the proportions.
3. The method for detecting corrosion of fasteners on overhead power transmission lines based on deep learning according to claim 2, characterized in that, In S2, the specific construction process of the multi-scale corrosion sensing enhancement feature extraction unit is as follows: The image of fastener corrosion in any power transmission corridor is denoted as fastener corrosion detection feature map F1. It is used as the input of this unit and passed through a CBS module consisting of a convolutional layer, a batch normalization layer and a SiLU activation function to output fastener corrosion detection feature map F2. Input F2 into the CBS module to obtain the fastener corrosion detection feature map F3; Inputting F3 into the context-aware branch will output fastener corrosion detection feature maps F5, F8, and F11 in sequence; simultaneously, inputting F3 into the detail-focusing branch will output fastener corrosion detection feature maps F6, F9, and F12 in sequence. Perform a Concat operation on F5 and F6, and stitch them together along the channel dimension to obtain the fastener corrosion detection feature map F7; then perform a Conv1×1 convolution operation on F7 to obtain the high-resolution corrosion detail feature map P1. Perform a Concat operation on F8 and F9, and concatenate them along the channel dimension to obtain the fastener corrosion detection feature map F10; then perform a Conv1×1 convolution operation on F10 to obtain the balanced corrosion structure feature map P2. Perform a Concat operation on F11 and F12 and concatenate them along the channel dimension to obtain the fastener corrosion detection feature map F13; then perform a Conv1×1 convolution operation on F13 to obtain the strong semantic corrosion diagnosis feature map P3.
4. The method for detecting corrosion of fasteners on overhead power transmission lines based on deep learning according to claim 3, characterized in that, In S3, the specific construction process of the cross-layer adaptive pyramid feature fusion unit is as follows: P1, P2, and P3 are input into the first CHMC submodule to obtain fastener corrosion detection feature map F14. Then, F14 is upsampled to obtain fastener corrosion detection feature map F15. Through the Concat operation, F15 and P1 are concatenated along the channel dimension to obtain fastener corrosion detection feature map F17. F17 is input into the C2F module for feature integration and transformation to obtain fastener corrosion detection feature map F18. At the same time, F14 is downsampled to obtain fastener corrosion detection feature map F16. Through the Concat operation, F16 and P3 are concatenated along the channel dimension to obtain fastener corrosion detection feature map F20. F20 is input into the C2F module for feature integration and transformation to obtain fastener corrosion detection feature map F21. Input F15 and F18 together into the first DIFA submodule to obtain fastener corrosion detection feature map F19; at the same time, input F16 and F21 together into the second DIFA submodule to obtain fastener corrosion detection feature map F22. F14, F18, and F21 are input into the second CHMC submodule to obtain the balanced collaborative corrosion correlation feature map P5. Then, P5 is upsampled to obtain the fastener corrosion detection feature map F23. Through the Concat operation, F23 and F19 are concatenated along the channel dimension to generate the fastener corrosion detection feature map F25. F25 is input into the C2F module for final feature purification and enhancement to obtain the detail-enhanced corrosion localization feature map P4. At the same time, P5 is downsampled to obtain the fastener corrosion detection feature map F24. Through the Concat operation, F24 and F22 are concatenated along the channel dimension to obtain the fastener corrosion detection feature map F26. F26 is input into the C2F module for final feature purification and enhancement to obtain the semantically driven corrosion discrimination feature map P6.
5. The method for detecting corrosion of fasteners on overhead power transmission lines based on deep learning according to claim 1, characterized in that, The specific steps of S4 are as follows: Based on the multi-scale corrosion perception enhancement feature extraction unit and the cross-layer adaptive pyramid feature fusion unit, an MS-AFRDN model is constructed for corrosion detection of overhead line fasteners in the power transmission corridor inspection scenario. The MS-AFRDN model is input into any corrosion image of a fastener in the power transmission corridor and finally outputs the structured information of the corrosion fastener, including corrosion status level, bounding box and confidence score. The MS-AFRDN model was trained based on a dataset of images of fastener corrosion in power transmission corridors.