A power transmission line icing fine segmentation method for complex background interference

By constructing a three-segment network structure with a well-cleaned and labeled dataset and a multi-scale feature enhancement module, the accuracy problem of icing region segmentation of transmission lines under complex backgrounds was solved, achieving high-precision pixel-level icing segmentation and improving the accuracy and robustness of icing detection.

CN122391628APending Publication Date: 2026-07-14NANJING UNIV OF INFORMATION SCI & TECH +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV OF INFORMATION SCI & TECH
Filing Date
2026-03-03
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing deep learning-based semantic segmentation methods struggle to stably segment icing areas of transmission lines under complex background interference, making it difficult to accurately identify the icing area and background edge contours, thus affecting the accuracy of icing thickness calculation.

Method used

A semantic segmentation model for icing on transmission lines is constructed, adopting a three-segment network structure of encoder-feature enhancement module-decoder. It combines the SegFormer semantic segmentation model architecture and the multi-scale feature enhancement U-shaped network MSFE-USN. By constructing a clean and labeled dataset and introducing a multi-scale feature enhancement module, the interference of complex backgrounds is suppressed and the fine-grained features of icing are enhanced.

Benefits of technology

Achieving high-precision pixel-level ice segmentation in complex backgrounds improves the accuracy and robustness of ice detection, providing reliable segmentation results to support subsequent ice thickness calculation and risk assessment, and adapting to different image quality conditions while maintaining stable segmentation performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391628A_ABST
    Figure CN122391628A_ABST
Patent Text Reader

Abstract

The application discloses a power transmission line icing fine segmentation method for complex background interference, comprising the following steps: constructing a pixel-level labeling data set containing a basic set and a generalization verification set; constructing a three-section segmentation model of an encoder-feature enhancement module-decoder, wherein the encoder and the decoder are based on a SegFormer architecture; introducing a multi-scale feature enhancement U-shaped network MSFE-USN as the feature enhancement module; the MSFE-USN effectively suppresses complex background interference and strengthens the edge and detail features of the icing area by integrating an attention mechanism and multi-scale feature fusion; after training and optimization of the model by using the training set, high-precision pixel-level icing segmentation of the power transmission line image under the complex background is realized; and the application significantly improves the segmentation accuracy of the model under the complex background, can accurately identify the edge contour between the icing area and the background, and provides more reliable segmentation result support for subsequent icing thickness calculation and risk assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent detection technology for icing on power transmission lines, and particularly relates to a refined segmentation method for icing on power transmission lines with complex background interference. Background Technology

[0002] With the sustained and rapid development of my country's economy, electricity demand has steadily increased, and power grid construction has entered a stage of rapid development. To optimize energy allocation and improve power supply reliability and stability, my country is actively promoting the coordinated integration of all aspects of the power grid, constructing a high-voltage, high-capacity, long-distance power transmission network system. However, the actual operating environment of transmission lines is complex and diverse, often crossing different geographical regions and climate zones. Especially under extreme weather conditions such as freezing rain, strong winds, heavy fog, and blizzards, coupled with factors such as micro-topography and micro-meteorology, safety accidents such as line icing, galloping, and even line breaks and tower collapses are easily triggered. Icing disasters have become one of the major threats facing the power grid system, and in severe cases, may lead to prolonged and large-scale paralysis of the transmission system, seriously affecting social production and daily life.

[0003] Currently, traditional techniques are mainly used for monitoring icing on transmission lines both domestically and internationally, including manual ice observation, simulated conductor methods, and fiber optic sensor methods. Manual ice observation relies on visual inspection by on-site personnel; simulated conductor methods estimate ice thickness through indirect mechanical analysis; and fiber optic sensor methods utilize distributed measurement for icing monitoring. With technological advancements, icing detection methods based on online monitoring equipment and image analysis technology have gradually become a research hotspot. This method uses cameras mounted on towers or drones for non-contact image acquisition, employs image processing techniques to segment the icing area, and then calculates the ice thickness based on pixel conversion relationships. In recent years, semantic segmentation methods based on deep learning have made phased progress in this field, achieving automatic identification and segmentation of icing areas by constructing convolutional neural network models.

[0004] While existing deep learning-based semantic segmentation methods have been applied in icing detection, they still have significant shortcomings in real-world complex environments. The most obvious drawback is the poor robustness of existing models to complex background interference, making it difficult to stably segment iced regions under multiple disturbances such as rain, snow, fog, changes in lighting, and obstruction by foreign objects. Furthermore, the blurred edges and low feature recognition between the iced region and the background make it difficult for the model to accurately identify the icing contour, directly affecting the accuracy of subsequent icing thickness calculations. These shortcomings limit the reliable application of existing methods in real-world complex scenarios and cannot meet the requirements for high-precision, high-robustness icing monitoring. Summary of the Invention

[0005] Purpose of the Invention: The purpose of this invention is to provide a refined segmentation method for icing on transmission lines in complex background interference. This method addresses the technical shortcomings of existing deep learning semantic segmentation methods in transmission line icing detection, such as unstable image quality and difficulty in accurately identifying the edge contours of icing areas and backgrounds in complex backgrounds. It improves the segmentation accuracy of icing areas on transmission lines in complex environments, providing reliable segmentation results to support subsequent key tasks such as icing thickness calculation and icing risk assessment.

[0006] Technical solution: The refined segmentation method for icing of transmission lines against complex background interference, as described in this invention, includes the following steps:

[0007] S1. Construct a semantic segmentation dataset for icing of transmission lines. Clean the original transmission line images, perform pixel-level annotation and dataset partitioning to form a basic dataset for model training and testing, and construct a generalization verification dataset across regions and devices.

[0008] S2. Based on the three-segment network structure of encoder-feature enhancement module-decoder, a semantic segmentation model for icing of transmission lines is built. The encoder and decoder are based on the SegFormer semantic segmentation model architecture and are used to extract and reconstruct multi-scale semantic features. The feature enhancement module is a multi-scale feature enhancement U-shaped network MSFE-USN, which is used to suppress complex background interference and enhance fine-grained features of icing.

[0009] S3. Train the transmission line icing semantic segmentation model using the training set in the basic dataset, and perform hyperparameter tuning and performance monitoring using the validation set in the basic dataset to obtain the trained transmission line icing semantic segmentation model.

[0010] S4. Input the transmission line image to be segmented, which contains a complex background, into the trained transmission line icing semantic segmentation model, and output the pixel-level semantic segmentation result.

[0011] This invention effectively improves the data quality and generalization ability of the model training by constructing a well-cleaned and labeled basic dataset containing cross-regional generalization validation data. Furthermore, it designs a three-segment network structure of encoder-feature enhancement-decoder, introducing a multi-scale feature enhancement U-shaped network MSFE-USN based on SegFormer multi-scale feature extraction and reconstruction. This significantly suppresses interference from complex backgrounds and enhances the fine-grained feature expression of icing. Through targeted training and hyperparameter tuning, the model can stably adapt to different image quality conditions and accurately identify the edge contours between the icing area and the background. Finally, it achieves high-precision pixel-level icing segmentation in complex backgrounds, providing reliable and robust segmentation results for subsequent tasks such as icing thickness calculation and risk assessment, thus improving the practicality and accuracy of transmission line icing detection in complex environments.

[0012] Preferably, the construction method of the basic dataset mentioned in step S1 includes:

[0013] Original power transmission line images covering different time periods, weather conditions, and geographical backgrounds were collected, and the original power transmission line images were standardized and cleaned to remove invalid samples.

[0014] The cleaned transmission line images were manually annotated at the pixel level, and three semantic categories were defined: background, main view transmission line, and side view transmission line.

[0015] The labeled transmission line images are randomly divided into training subsets, validation subsets, and basic test subsets according to a preset ratio, and each subset is ensured to maintain a balanced distribution in terms of time, weather, background, and icing status.

[0016] By collecting diverse images covering different time periods, weather conditions, and geographical backgrounds, and performing standardized cleaning, data quality and broad representativeness were ensured from the source. Then, through manual pixel-level fine annotation and clear differentiation into three categories—background, main-view transmission lines, and side-view transmission lines—clear and accurate learning objectives were provided for the model, especially for lines observed from different angles. Finally, by ensuring a balanced distribution of each data subset across multiple dimensions, model overfitting or performance imbalance caused by data bias was effectively avoided. This laid a high-quality, highly generalizable, and class-balanced data foundation for subsequent model training and evaluation, significantly improving the model's adaptability to complex and ever-changing real-world scenarios.

[0017] Preferably, the generalization validation dataset described in step S1 is constructed in the following ways:

[0018] Collect images of power transmission lines that differ from the base dataset in terms of camera model, geographic background type, and icing physical type.

[0019] The collected transmission line images with differences are cleaned and labeled at the pixel level, and the labeling specifications are consistent with the basic dataset.

[0020] All labeled transmission line images are used as an independent generalization validation test set, which is not included in the model training process, and is used to evaluate the model's cross-domain generalization ability.

[0021] By constructing an independent test set that differs systematically from the basic training set in terms of shooting equipment, geographical background, and icing physics, and strictly adhering to the same cleaning and labeling standards as the basic dataset, this step creates an evaluation benchmark that can rigorously and objectively test the model's cross-domain adaptability. This dataset is not used in the training process and is specifically designed to simulate and evaluate the model's robustness and generalization performance when faced with unfamiliar equipment models, unknown geographical environments, and unseen icing morphologies. This effectively avoids the risk of overfitting the model on familiar data and ensures that the developed icing segmentation method has reliable transfer capabilities and stable performance in complex and ever-changing application scenarios. This provides crucial validation for the model's universality and reliability in actual engineering deployments.

[0022] Preferably, the multi-scale feature enhancement U-shaped network MSFE-USN described in step S2 adopts a symmetrical U-shaped structure, which includes a top-down upsampling path and several lateral skip connection paths;

[0023] The encoder outputs four layers of feature maps, denoted as F1, F2, F3, and F4, with their resolution decreasing sequentially. MSFE-USN processes these feature maps using the following process:

[0024] S21. Deep feature processing: The deepest feature map F4 is input into a convolutional attention module CBAM for preliminary calibration, and then the calibrated feature map is input into an edge-guided multi-scale attention module EMSAM to output an enhanced feature map F5 that combines high-level semantics and rich edge details.

[0025] S22. Feature Fusion and Upsampling: Upsample the enhanced feature map F5 by 2x bilinear interpolation to match the resolution of feature map F3, and then fuse it with the shallow feature map F3 enhanced by CBAM; use 1×1 convolution to unify the number of channels on the fusion result, and repeat the steps of upsampling, feature fusion with the shallow feature maps F2 and F1 enhanced by CBAM, and channel unification.

[0026] The calculation method for feature fusion is as follows:

[0027]

[0028]

[0029] in, This represents the feature map of the i-th layer output by the encoder. This represents the intermediate enhanced feature map obtained during the stepwise upsampling and fusion process of MSFE-USN, used to carry the fusion result of deep semantic information and corresponding shallow detail information. This indicates a 2x bilinear interpolation upsampling. This represents a 1×1 convolution operator. This represents the convolutional attention module. This indicates an edge-guided multi-scale attention module;

[0030] S23. Feature Output: Output a set of enhanced multi-scale feature maps and provide them to the decoder for final segmentation prediction.

[0031] The multi-scale feature enhancement U-shaped network MSFE-USN constructs an enhancement path that deeply integrates shallow details and deep semantics through its symmetrical U-shaped structure and lateral jump connections. First, it uses the convolutional attention module CBAM and the edge-guided multi-scale attention module EMSAM to perform dual calibration on the deep feature map, significantly enhancing its ability to focus on high-level semantic information of the icy region, and especially improving its sensitivity and representation ability for complex edge contours. Subsequently, through bilinear upsampling, fusion with attention-enhanced shallow features, and channel unification operations, it systematically fuses refined deep semantic information with rich shallow spatial details at multiple levels, effectively suppressing background noise interference, while accurately recovering the fine structure and edge information of the icy target at multiple scales. Finally, the output set of enhanced multi-scale feature maps provides the decoder with high-quality feature representation that combines global contextual understanding and local detail resolution, thus laying a key core feature foundation for the model to achieve high-precision and robust pixel-level icing segmentation in complex backgrounds.

[0032] Preferably, the convolutional attention module CBAM processes the input feature map sequentially through a channel attention submodule and a spatial attention submodule;

[0033] The channel attention submodule first performs global average pooling and global max pooling on the input feature map simultaneously to obtain two one-dimensional feature vectors; the two one-dimensional feature vectors are then fed into a shared two-layer multilayer perceptron for processing; the two processed one-dimensional feature vectors are added element by element, and then a channel attention weight vector is generated by passing the sigmoid activation function.

[0034] The spatial attention submodule receives the feature map calibrated by the channel attention weight vector as input. First, it performs global average pooling and global max pooling along the channel dimension to obtain two two-dimensional feature maps. The two two-dimensional feature maps are then concatenated along the channel dimension. The concatenated two-dimensional feature map is then convolved through a 7×7 convolutional layer and then activated by the Sigmoid activation function to generate a spatial attention weight map.

[0035] The Convolutional Attention Module (CBAM) achieves dual-dimensional adaptive calibration of the input feature map from feature channels to spatial location through a cascaded channel and spatial attention mechanism. Its channel attention submodule, by fusing complementary information extracted by global average pooling and max pooling and using a shared multilayer perceptron for nonlinear transformation, can accurately learn and enhance feature channels that are crucial for icing segmentation, effectively suppressing interference from irrelevant background channels. Subsequently, the spatial attention submodule, based on channel enhancement, further fuses spatial context information obtained from different pooling operations and captures extensive neighborhood relationships through large-size convolution, thereby generating an attention weight map that can accurately focus on the spatial location and contour of the icing area. The two work together to enable the model to dynamically and discriminatively enhance key features related to transmission line icing while weakening the response to complex background areas, significantly improving the discriminative power and robustness of feature representation.

[0036] Preferably, the edge-guided multi-scale attention module (EMSAM) is composed of the following cascaded units:

[0037] Feature extraction and fusion unit: It contains a convolutional module for extracting input features; then, max pooling and average pooling operations are performed in parallel on the extracted features, and the two pooling results are concatenated along the channel dimension for preliminary fusion of local details and global context information;

[0038] Feature enhancement and filtering unit: Connected after the feature extraction and fusion unit, it processes the concatenated features through a convolutional layer, and then generates a feature weight map through the Sigmoid activation function; the feature weight map is then multiplied element-wise with the original convolutional output of the feature extraction and fusion unit to perform preliminary feature enhancement and noise filtering;

[0039] Multi-scale context aggregation unit: connected after the feature enhancement and filtering unit, it is a dilated spatial convolutional pooling pyramid (ASPP) structure; the ASPP structure contains at least three parallel branches, each using a 3×3 dilated convolutional layer with different dilation rates, to simultaneously capture global and local contextual information under different receptive fields, and to enhance edge feature responses at multiple scales.

[0040] Two-dimensional fine calibration unit: connected after the multi-scale context aggregation unit, it is a parallel spatial and channel squeezing excitation scSE module; the scSE module runs a channel attention submodule cSE and a spatial attention submodule sSE in parallel, and adds the output feature maps after calibration of the two submodules along the channel direction, which is used to calibrate the spatial dimension and channel dimension of the feature map synchronously.

[0041] The edge-guided multi-scale attention module (EMSAM) achieves multi-level edge feature enhancement and calibration from local details to global context through cascaded units. The feature extraction and fusion unit initially integrates local details and global context information through parallel pooling operations. Subsequently, the feature enhancement and filtering unit effectively filters and enhances key edge-related features and suppresses irrelevant noise by multiplying features with adaptive weight generation. Next, the multi-scale context aggregation unit (ASPP) uses dilated convolutions with different dilation rates to simultaneously capture multi-scale context information from local fine structure to global semantic relationships, significantly enhancing the ability to represent complex edge contours at different scales. Finally, the dual-dimensional fine calibration unit (scSE) implements attention mechanisms in parallel across spatial and channel dimensions, performing synchronous fine calibration of the feature map's spatial location and feature channels to ensure accurate focus and response to icing edge features. Through this series of progressively layered processes, the module enables the model to keenly perceive and enhance the fine-grained edge features of icing on transmission lines, thereby achieving more accurate edge segmentation in complex backgrounds.

[0042] Preferably, the multi-scale context aggregation unit further includes an image pooling branch, which sequentially performs global average pooling, 1×1 convolution, and upsampling operations.

[0043] The introduction of this image pooling branch captures the broadest global contextual information covering the entire input feature map through global average pooling operations and compresses it into a highly abstract global feature representation. Subsequently, feature transformation and dimensionality adaptation are performed through 1×1 convolution, and spatial resolution is restored through upsampling operations. Finally, the contextual information containing image-level global semantics is effectively fused into the multi-scale feature stream. This mechanism enables the multi-scale context aggregation unit to not only obtain multi-scale local context in parallel through dilated convolutions with different dilation rates, but also to supplement the most generalized global scene understanding. This enhances the model's ability to respond to local edge details while ensuring that its decision-making process is effectively guided by the overall scene semantics, significantly improving the model's recognition accuracy and segmentation consistency of the overall structure and distribution of icy targets in complex backgrounds.

[0044] Preferably, in the parallel spatial and channel compression excitation scSE module of the dual-dimensional fine calibration unit, the channel attention submodule cSE generates a channel attention weight vector by performing global average pooling on the input feature map and then processing it through a two-layer multilayer perceptron and a sigmoid function; the spatial attention submodule sSE generates a spatial attention weight map by performing 1×1 convolution on the input feature map and then processing it through a sigmoid function.

[0045] The parallel scSE module in this dual-dimensional fine-tuning calibration unit achieves simultaneous fine-tuning of feature map channels and spatial dimensions through the synergistic effect of the channel attention submodule cSE and the spatial attention submodule sSE. The cSE submodule captures global statistical information of feature channels through global average pooling and learns the nonlinear dependencies between channels through a multilayer perceptron, thereby generating attention weights that can dynamically enhance task-relevant feature channels and suppress irrelevant channels, effectively optimizing the channel representation capability of the feature map. At the same time, the sSE submodule efficiently learns and generates attention weight maps in the spatial dimension through lightweight 1×1 convolution operations, which can accurately focus on key spatial locations in the feature map related to the icing area and edges, and suppress interference from the background area. The outputs of the two submodules are fused by addition, so that the final calibrated feature map is adaptively enhanced in terms of channel importance and spatial saliency, significantly improving the model's ability to distinguish subtle features of icing targets and segmentation accuracy.

[0046] Preferably, in the lateral skip connection path, the multi-scale feature enhancement U-shaped network MSFE-USN adopts a collaborative enhancement strategy of CBAM and EMSAM cascade for the deepest feature map link, including: on the lateral connection path corresponding to the deepest feature map F4, the deepest feature map F4 is first preliminarily calibrated in terms of channel and spatial dimensions through the CBAM module to suppress invalid noise and highlight task-related regions; then the output of the CBAM module is used as the input of the edge-guided multi-scale attention module EMSAM, and multi-scale dilated convolution and parallel spatial channel attention are executed synchronously through EMSAM to aggregate global and local context and enhance the details of icing edges, forming a cascaded processing flow of first filtering and denoising, and then edge enhancement.

[0047] In the deepest lateral connection path of MSFE-USN, a collaborative enhancement strategy of cascading CBAM and EMSAM is adopted to construct a refined feature processing flow that first filters and denoises, and then enhances edges. First, the CBAM module performs preliminary calibration of the deepest feature map F4 in terms of channel and spatial dimensions, effectively suppressing invalid feature responses caused by background clutter and noise interference, and initially focusing on key semantic regions related to icing. Subsequently, the purified and focused features are input into the EMSAM module, which, through its multi-scale dilated convolutional structure and parallel spatial channel attention mechanism, simultaneously aggregates multi-level contextual information from local details to global semantics, and significantly enhances the fine-grained feature representation capability of icing regions, especially their complex edge contours. This cascaded design realizes a progressive enhancement by first extracting high-purity semantic information from deep features and then injecting multi-scale edge details, so that the features finally fused into the decoding path have both clear semantic guidance and rich edge information, providing a strong discriminative deep feature foundation for the core segmentation task.

[0048] Preferably, the decoder described in step S2 is used to perform channel unification, upsampling and fusion processing on the multi-scale features output by the feature enhancement module to generate pixel-level probability maps corresponding to each semantic category, and obtain the final pixel-level semantic segmentation result through normalization processing.

[0049] The decoder systematically unifies the channel dimensions and upsamples the spatial resolution of the multi-scale features output by the feature enhancement module, effectively integrating and reconstructing deep semantic information and shallow detail information at different scales. This generates clear, coherent pixel-level probability maps corresponding to each semantic category. This process ensures a smooth transition from rich feature representation to accurate spatial localization. In particular, by fusing multi-scale enhanced features, it significantly improves the model's ability to distinguish foreground targets (such as main-view and side-view transmission lines) from complex backgrounds and to finely depict the edge contours of icing areas. Finally, through normalization, the model can output stable and reliable pixel-level semantic segmentation results, providing a high-precision and highly consistent segmentation foundation for subsequent quantitative analysis tasks such as accurate calculation of icing thickness and risk assessment.

[0050] Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: 1. By constructing a three-segment network architecture of encoder-feature enhancement module-decoder, this invention designs a multi-scale feature enhancement module MSFE-USN that can suppress background interference and enhance fine-grained features of icing. In the feature extraction and fusion process, it effectively enhances the modeling ability of edge details and multi-scale contextual information, significantly improves the segmentation accuracy of the model in complex backgrounds, and can accurately identify the edge contours between the icing area and the background, providing more reliable segmentation results for subsequent icing thickness calculation and risk assessment; 2. By introducing an attention mechanism in the multi-scale feature enhancement module, this invention can adaptively focus on the icing area of ​​the transmission line, effectively suppressing noise interference from diverse and complex backgrounds such as rain, snow, fog, haze, forests, and lake reflections. This design enables the model to maintain stable and accurate segmentation performance even in extreme environments with unstable image quality or visual confusion; 3. By constructing and utilizing training and validation datasets covering different regions, equipment, and environmental conditions, the model learns more universal feature representations. Cross-domain validation showed that the model maintained high segmentation accuracy even when faced with transmission line images taken by different devices or in different geographical backgrounds outside the training data distribution, demonstrating strong generalization performance and adapting to the diverse monitoring scenarios in practical applications. 4. While improving segmentation performance, this invention also features a carefully designed model structure, achieving a good balance between computational complexity and the number of parameters. The model maintains a high inference speed, enabling it to meet the real-time or near-real-time processing requirements of practical deployments, providing a feasible technical solution for online icing monitoring on edge devices or mobile platforms with limited computing power. Attached Figure Description

[0051] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0052] Figure 2 This is a schematic diagram of the model structure of the present invention;

[0053] Figure 3 This is a diagram of the multi-scale feature enhancement U-shaped network structure of the present invention;

[0054] Figure 4 This is a schematic diagram of the edge-guided multi-scale attention module of the present invention;

[0055] Figure 5 This is a structural diagram of the CBAM and its different modules according to the present invention;

[0056] Figure 6 A comparison diagram of the icing segmentation results of transmission lines under different complex and extreme environments according to the present invention;

[0057] Figure 7This is a comparison diagram of the refined icing segmentation results of the present invention. Detailed Implementation

[0058] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0059] 1. Implementation prerequisites and environmental preparation

[0060] 1.1 Hardware Environment: The hardware configuration used in this implementation is as follows to ensure efficient operation of model training and inference: CPU: Intel Core i9-12900K (16 cores, 24 threads, 3.2GHz, 30MB cache); GPU: NVIDIA RTX 3090 (24GB VRAM, 10496 CUDA cores); Memory: 64GB DDR5 4800MHz; Storage: 2TB SSD for storing datasets, model weights, and training logs.

[0061] 1.2 Software Environment: The operating system used is Ubuntu 20.04 LTS 64-bit; the deep learning framework is PyTorch 1.12.1, paired with CUDA 11.6 and CuDNN 8.4.1 to accelerate GPU computing; the programming language is Python 3.8.13; the dependent libraries include: OpenCV 4.6.0 (image reading and preprocessing), LabelMe 4.6.0 (dataset labeling), NumPy 1.23.5 (numerical computation), Matplotlib 3.6.2 (result visualization), Scikit-learn 1.1.3 (data partitioning and index calculation), and tqdm 4.64.1 (training progress display).

[0062] 2. Construction and Implementation of Transmission Line Icing Dataset

[0063] Figure 1 This is a flowchart of the method in this embodiment. This implementation method constructs two types of datasets (basic dataset and generalization validation dataset), and the specific implementation steps are as follows:

[0064] 2.1 Data Acquisition: The original images of the basic dataset are from the icing monitoring site of transmission lines in Anhui Province. They were collected by fixed cameras on the towers, covering multiple time periods such as early morning, daytime, dusk, and night, and encompassing meteorological conditions such as sunny, cloudy, rainy, snowy, and foggy days. The backgrounds include complex scenes such as valleys, lakes, roads, and towers. The original images of the generalization validation dataset are from the monitoring data of transmission lines in Shanxi Province provided by Shanxi Electric Power Research Institute. The backgrounds are mainly plains and hills, and the icing type is mainly rime.

[0065] 2.2 Data Cleaning: The raw images were processed using a standardized process: ① Unified naming rules ② Invalid sample removal: Through manual screening combined with image clarity assessment, samples that were severely blurred, had excessively large target occlusion areas, or could not be identified due to icing were removed. In the end, 5140 valid raw images (Anhui region) and 806 valid raw images (Shanxi region) were retained.

[0066] 2.3 Pixel-level annotation: Manual annotation was performed using the LabelMe annotation tool. Annotators, after professional training, completed the annotation according to the preset semantic categories (background, main view transmission line, side view transmission line): ① For single-conductor images, the main view transmission line was directly annotated; ② For multi-conductor images, the clearest conductor that best met the segmentation requirements was selected as the main view transmission line, and the remaining significant conductors were designated as side view transmission lines. During annotation, the outline was accurately drawn along the edge of the conductor to ensure complete coverage of the icing area; ③ After annotation, the data was exported as a JSON file and converted into a VOC format label image using a custom script.

[0067] 2.4 Dataset Partitioning: 5140 labeled images from Anhui Province were randomly divided into training, validation, and test sets in a 6:2:2 ratio, with 3084 images in the training set, 1028 in the validation set, and 1028 in the test set. To ensure the stability of model training and the reliability of validation and testing results, stratified sampling was used to ensure a balanced distribution of each subset. Statistical analysis revealed clear characteristics in the original dataset across multiple dimensions: in the time dimension, daytime and nighttime samples were approximately 1:1; in the weather dimension, sunny, cloudy, foggy, and rainy / snowy days were approximately 2:2:3:3; and in the target state dimension, icy and non-iced samples were approximately 5:1. Based on this, the training, validation, and test sets maintained the same distribution ratios as the original dataset in the core dimensions of time distribution, weather type, background scene, and icing status, achieving a roughly balanced distribution across the subsets. To verify the generalization performance of the proposed model, 806 labeled images from Shanxi Province were used as a dedicated test set for generalization verification, with the ratio of iced to uniced samples being approximately 6:1.

[0068] 2.5 Data Augmentation: To improve the model's generalization ability, online augmentation was performed on the training set images. Augmentation strategies included: ① random flipping; ② random cropping; ③ brightness and contrast adjustment; ④ Gaussian noise addition; ⑤ random scaling. The validation and test sets were not augmented, only their sizes were normalized.

[0069] 3. Improve the construction and implementation of the SegFormer model.

[0070] The embodiments of this invention are based on SegFormer-B2 to build an improved model. The specific building steps are as follows:

[0071] 3.1 Basic Network Loading: Download the pre-trained SegFormer-B2 model weights from the PyTorch official model library and load its encoder and decoder structures.

[0072] 3.2 MSFE-USN Module Setup: An MSFE-USN module is connected in series between the encoder and decoder. The specific structure is as follows: Figure 2 As shown:

[0073] 3.2.1 U-shaped architecture construction: A symmetrical U-shaped structure is adopted, which includes a top-down feature transmission path and a lateral jump connection path: ① Top-down path: The F4 feature map is upsampled by 2 times (bilinear interpolation) to obtain feature maps that match the resolution of F3, F2 and F1; ② Lateral jump connection path: The encoder F1-F4 and the corresponding feature maps of the top-down path are connected respectively to form 4 jump connection branches.

[0074] 3.2.2 CBAM Module Embedding: Embed a CBAM module in each lateral jump connection branch, such as... Figure 5 As shown, the parameters of the CBAM module are set as follows: ① Channel attention submodule: global average pooling and global max pooling are used for joint representation, the number of MLP layers is 2, the number of input channels is 128, the number of hidden layer channels is 16 (reduction rate r=8), and the activation function is ReLU; ② Spatial attention submodule: a 7×7 convolution kernel is used, the number of input channels is 2, the number of output channels is 1, and the activation function is Sigmoid; the input of the CBAM module is the encoder output feature map, and the output is the feature map after channel and spatial calibration.

[0075] 3.2.3 EMSAM Module Embedding: In the lateral jump connection branch corresponding to F4, an EMSAM module is connected in series after the CBAM module, as shown in the specific structure. Figure 3As shown, its structure is as follows: ① Feature extraction and initial fusion: After the input features are initially extracted by the Conv module, max pooling and average pooling operations are performed simultaneously. The two pooling results are concatenated along the channel dimension to achieve initial fusion of local detail features and global context information; ② Feature enhancement and noise filtering: The concatenated features are dimensionality reduced and integrated by the convolutional layer, and then a feature weight map is generated by the Sigmoid activation function. This weight map is multiplied element-wise with the original feature map output by the Conv module to complete feature enhancement and initial suppression of invalid noise. The processed features are then input into the ASPP module; ③ Multi-scale context aggregation: The ASPP module achieves multi-scale feature capture through parallel branch design, specifically including three 3×3 dilated convolution branches with different dilation rates (6, 12, 18) and one image pooling branch (global average pooling + 1×1 convolution + upsampling). Through parallel operation of multiple branches, feature information under different receptive fields is captured simultaneously to achieve deep aggregation of global and local features, significantly improving the model's ability to perceive subtle edge structures in complex backgrounds; ④ Feature fine-tuning calibration: such as Figure 4 As shown, the output of the ASPP module is directly input into the scSE module for accurate calibration. The scSE module is an improvement and optimization of the traditional SE attention module, including three functional variants: cSE, sSE, and scSE. Their core function is to enhance features relevant to the icing segmentation task and suppress irrelevant background features. The cSE module introduces a channel attention mechanism, using global average pooling and an MLP network to adaptively learn channel weights, thereby integrating and enhancing channel-dimensional feature information. The sSE module focuses on spatial dimension optimization, first extracting spatial weight information through 1×1 convolution to generate a spatial attention map, and then multiplying it element-wise with the original feature map to achieve spatial dimension attention enhancement. The scSE module is a collaborative combination of the sSE and cSE modules, simultaneously optimizing both the spatial and channel dimensions of the feature map. Finally, the outputs of the two modules are added along the channel direction to achieve a more comprehensive and accurate feature enhancement effect.

[0076] 3.3 Decoder and Output Layer Configuration: The lightweight MLP decoder of SegFormer-B2 is used. The input is the enhanced feature map output by MSFE-USN. The decoder unifies the number of channels of each enhanced feature map through 1×1 convolution, then upsamples through bilinear interpolation, concatenates and fuses the features through 3×3 convolution. Finally, it outputs a 3-channel feature map (corresponding to 3 semantic categories) through 1×1 convolution. The class probability distribution of each pixel is obtained through the Softmax activation function.

[0077] 4. Implementation of Model Training and Validation

[0078] 4.1 Training Parameters and Performance Evaluation Metrics: The core hyperparameters for model training in this implementation are configured as follows: the Adam optimizer is used, the initial learning rate is set to 0.0001, the batch size is set to 32, and the number of training epochs is set to 100. To objectively and comprehensively evaluate the model's performance in the transmission line icing segmentation task, this study adopts the mainstream evaluation metric system in the semantic segmentation field. The core metrics include Mean Intersection over Union (mIoU), Mean Pixel Accuracy (mPA), and Mean Precision (mPrecision). All evaluation metrics are calculated as averages across all semantic categories: "background, main-view transmission line, and side-view transmission line," ensuring a balanced reflection of the model's segmentation capabilities for different categories and avoiding evaluation bias caused by differences in the distribution of samples from a single category, thus guaranteeing the objectivity and reliability of the evaluation results.

[0079] 4.2 Training Process Implementation: ① Data Loading: Load the training and validation sets using a custom DataLoader, enabling multi-threaded reading (num_workers=8) to achieve online data augmentation and batch loading; ② Model Initialization: Move the completed improved model to the GPU device and initialize the parameters of the newly added modules (MSFE-USN, EMSAM) (using Xavier initialization); ③ Forward Propagation: Input training set images, which are processed by the encoder, MSFE-USN, and decoder to obtain the predicted probability map; ④ Loss Calculation: Calculate the loss between the predicted probability map and the label map, backpropagate the gradient, and update the model parameters; ⑤ Validation Process: After each round of training, calculate metrics such as mIoU, mPA, and mPrecision on the validation set, and save the model with the highest mIoU on the validation set as the optimal model. The training log records the loss value and evaluation metrics in real time.

[0080] 4.3 Model Validation Implementation:

[0081] (1) Ablation experiment verification: According to the combination shown in Table 1, five types of models were built in sequence: “SegFormer-B2 (baseline)”, “baseline + USN”, “baseline + USN + CBAM”, “baseline + USN + EMSAM”, and “baseline + USN + CBAM + EMSAM (model of this invention)”. After training with the same training parameters, the performance was evaluated on the test set in Anhui Province to verify the effectiveness of each module. Table 1 Ablation Experiment

[0082] Baseline USN CBAM EMSAM mIoU / (%) mPA / (%) mPrecision / (%) √ 92.91 96.68 95.78 √ √ 93.45 97.11 95.99 √ √ √ 93.81 97.25 96.24 √ √ √ 93.93 97.23 96.39 √ √ √ √ 94.07 97.37 96.40

[0083] (2) Comparative Experiment Verification: Fourteen mainstream models were selected, including the SegFormer series (B0-B5), DeepLab series (DeeplabV3plus-mobile, DeeplabV3plus-xception), U-Net series (Unet-vgg, Unet-resnet50), HRNet series (HRNetV2_w18, HRNetV2_w32, HRNetV2_w48), and PSPNet series (PSPNet-resnet50, PSPNet-mobilenet). After training with the same dataset and training parameters, the mIoU, mPA, and mPrecision metrics were evaluated on the test sets in Anhui and Shanxi provinces, respectively, to compare the performance advantages of the model proposed in this invention. The specific experimental results are as follows: The experimental results on the test set in Anhui province are shown in Table 2. The model proposed in this invention performed best among all the comparison methods, and the core metrics all reached the optimal level: mIoU was 94.07%, mPA was 97.37%, and mPrecision was 96.4%. Among the SegFormer series models of different sizes, SegFormer-B2 exhibits the best overall performance in the transmission line icing segmentation task, with an mIoU of 92.91%, mPA of 96.68%, and mPrecision of 95.78%, significantly outperforming SegFormer-B0, B1, B3, B4, and B5 versions. Compared to other versions, SegFormer-B2 has a moderate number of parameters and receptive field size, which can fully learn the detailed features of transmission line icing while avoiding overfitting on limited datasets due to excessive model size. Therefore, this invention selects SegFormer-B2 as the baseline model for improvement. The improved model's mIoU is improved to 94.07%, an improvement of 1.16 percentage points compared to the baseline model, fully validating the effectiveness of the improvement strategy of this invention. Among other mainstream models, DeeplabV3plus-xception and HRNetV2_w18 performed poorly in complex icing scenarios due to receptive field limitations and loss of detail features, with mIoU below 90%. The U-Net series, with its symmetric encoder-decoder structure, performed well in edge detail segmentation, with Unet-resnet50 achieving an mIoU of 92.65%, but still lower than the model presented in this invention. HRNetV2_w48 showed a significant performance drop, with an mIoU of only 68.69%, presumably due to overfitting on a limited dataset caused by excessive parameters. The PSPNet series also generally performed worse than the model presented in this invention due to insufficient capture of fine-grained edge features by multi-scale pooling operations. To verify the generalization performance of the proposed model, comparative experiments were conducted on a transmission line icing image dataset from Shanxi Province, and the results are shown in Table 3.The model of this invention maintains its leading overall performance on the Shanxi dataset, with mIoU of 93.11%, mPA of 96.31%, and mPrecision of 96.33%. Although slightly lower than the Anhui dataset, its overall performance is stable, fully demonstrating good generalization ability. Among the SegFormer series, SegFormer-B2 still performs best, with mIoU of 90.43%, mPA of 94.46%, and mPrecision of 94.95%. SegFormer-B0 and B1, due to their small model size and limited receptive field, struggle to capture both detailed ice-covering features and global correlation information simultaneously. SegFormer-B3 to B5, on the other hand, suffer from excessively large parameter counts, and their deep network structure tends to weaken edge features, thus limiting performance improvement. Compared to other mainstream models, DeeplabV3plus-xception and HRNetV2_w18, limited by their receptive field, performed poorly on the segmentation dataset of icing transmission lines in Shanxi Province. Unet-resnet50 ranked among the top models in comparison with an mIoU of 89.31%, but still lagged behind the model of this invention. HRNetV2_w48's mIoU on the Shanxi dataset was only 64.13%, further verifying its problem of overfitting on limited datasets due to its large number of parameters. The PSPNet series, on the other hand, continued to have low performance due to insufficient edge feature capture by multi-scale pooling.

[0084] Table 2 Comparison Experiment of Icing Datasets for Transmission Lines in Anhui Province

[0085] Model mIoU / (%) mPA / (%) mPrecision / (%) SegFormer-B0 92.55 96.62 95.45 SegFormer-B1 92.32 96.44 95.36 SegFormer-B2 92.91 96.68 95.78 SegFormer-B3 89.26 93.9 94.33 SegFormer-B4 89.7 94.35 94.41 SegFormer-B5 90.61 94.96 94.86 DeeplabV3plus-mobile 91.28 95.25 95.33 DeeplabV3plus-xception 89.54 94.9 93.61 Unet-vgg 91.96 95.97 95.41 Unet-resnet50 92.65 96.1 96.07 HRNetV2_w18 89.05 93.35 94.62 HRNetV2_w32 89.92 94.61 94.38 HRNetV2_w48 68.69 73.21 90.05 PSPNet-resnet50 87.36 92.09 93.87 PSPNet-mobilet 84.04 89.92 91.77 Our Model 94.07 97.37 96.4

[0086] Table 3 Comparative Experiment of Icing Datasets for Transmission Lines in Shanxi Province

[0087] Model mIoU / (%) mPA / (%) mPrecision / (%) SegFormer-B0 88.51 93.82 93.27 SegFormer-B1 89.15 93.99 93.87 SegFormer-B2 90.43 94.46 94.95 SegFormer-B3 81.63 87.71 90.04 SegFormer-B4 82.69 88.29 91.16 SegFormer-B5 82.05 86.34 92.55 DeeplabV3plus-mobile 86.67 90.52 94.19 DeeplabV3plus-xception 85.38 90.92 91.74 Unet-vgg 86.67 91.89 92.99 Unet-resnet50 89.31 93.04 95.06 HRNetV2_w18 84.55 88.04 94.14 HRNetV2_w32 85.69 89.11 94.47 HRNetV2_w48 64.13 67.78 83.99 PSPNet-resnet50 80.95 85.21 92.04 PSPNet-mobilet 77.67 82.18 90.19 Our Model 93.11 96.31 96.33

[0088] (3) Visual verification: through Figure 6 , Figure 7 Two types of visualization results comprehensively validate the model's performance, among which Figure 6 Focusing on the adaptability verification of complex and extreme scenarios, the segmentation results of various models in scenarios such as snow-covered valley roads, lake reflections, nighttime rain and snow, heavy fog, and overlapping solar panels on the Shanxi Plain show that existing mainstream models such as PSPNet-resnet50 and HRNetV2_w32 generally have problems such as background misjudgment (e.g., snow, lake surface, and rain and snow scratches are misjudged as ice), segmentation breakage, and blurred boundaries. They are not adaptable to scenarios with high similarity backgrounds, low texture interference, and low visibility. Figure 7Focusing on the performance verification of fine-grained segmentation, a comparison of the results of different models for fine-grained segmentation of icing on transmission lines shows that there are significant performance differences among the models in the fine-grained segmentation task: PSPNet-resnet50 performs the worst, failing to accurately capture the features of the icing area and being easily affected by background interference, for example in... Figure 7 (d) When processing the icy area of ​​the guide wire from the side viewpoint, the icing is incorrectly segmented together with the surrounding background, failing to effectively distinguish the target from the background. HRNetV2_w32 and Unet-resnet50 have similar problems; they fail to balance the information of icing and background at different scales during multi-scale feature fusion, resulting in limited segmentation accuracy. Furthermore, HRNetV2_w32... Figure 7 (c) exhibits significant segmentation breaks, failing to achieve continuous segmentation of the icy region and severely impacting the integrity of the segmentation results, reflecting its inadequacy in handling long-distance dependencies. SegFormer-B2, on the other hand, suffers from insufficient detail in local region feature representation and a lack of ability to capture detailed information, resulting in local breaks, incorrect segmentation, and imprecise differentiation between icing and background, indicating its insufficient performance in fine-grained icing segmentation in complex scenes. In contrast, the model proposed in this study effectively addresses these issues. Figure 6 , Figure 7 It performs optimally in various scenarios, effectively suppressing background interference and accurately distinguishing between icing and complex backgrounds. It also achieves continuous, complete, and high-precision icing segmentation. Furthermore, it demonstrates stronger generalization in unfamiliar cross-regional scenarios in Shanxi, which are not covered by the Anhui training set. Only in extreme scenarios where the background and conductor visual features are highly similar and spatially overlap, there is a slight fluctuation in the segmentation of conductor ends from the side view. This fully verifies the effectiveness, advancement, and completeness of the model in the transmission line icing segmentation task.

[0089] (4) Complexity Assessment: Under the same hardware environment, the FLOPS (computational cost) and Params (parameter count) of each model were statistically analyzed using the PyTorch Profiler tool, and the inference speed (FPS) was tested to assess the deployment feasibility of the model of this invention. Detailed data on the complexity of each model are shown in Table 4. The specific analysis is as follows: DeeplabV3plus-xception has a computational cost of 166.849G, a parameter count of 54.709M, and an FPS of 32.18. It has a complex structure and high resource requirements, which is not conducive to deployment on devices with limited computing power; DeeplabV3plus-mobile, on the other hand, has a computational cost of 52.875G, a parameter count of 5.814M, and an inference speed of 79.62FPS, showing advantages in lightweight and real-time performance, and is more suitable for real-time detection scenarios; Unet-resnet50 has a computational cost of 184.133G, a parameter count of 43.933M, and an FPS of 32.18. The computational cost of Unet-vgg is 33.06, which is structurally balanced but still has a relatively high computational burden. Unet-vgg has a computational cost of 451.706G and 24.891M parameters, with an FPS of only 17.51. Although the number of parameters is not significantly high, its real-time performance is poor. The number of parameters in the HRNet series models increases rapidly with the increase of network width, which leads to limited real-time performance and makes it difficult to adapt to mobile monitoring equipment. PSPNet-resnet50 has a computational cost of 118.447G, 46.716M parameters, and an FPS of 54.19, which is relatively balanced between speed and accuracy. PSPNet-mobilenet is extremely lightweight with an FPS of up to 121.64, which has outstanding potential in real-time scenarios. SegFormer-B2 has a computational cost of 113.452G, 27.349M parameters, and an FPS of 28.56, which achieves a certain balance between accuracy and efficiency. Finally, the model of this invention achieves a good trade-off between computational efficiency, parameter scale and inference speed with a computational load of 46.541G, a parameter count of 27.370M and an inference speed of 33.66FPS. It can maintain high availability performance in most practical monitoring scenarios, but there is still room for improvement in inference tasks with extremely high real-time requirements.

[0090] Table 4. Computational complexity results of each model in the task of fine segmentation of transmission line icing.

[0091] Model FLOPS Params FPS segformer-b2 113.452G 27.349M 28.56 unet-vgg 451.706G 24.891M 17.51 unet-resnet50 184.133G 43.933M 33.06 deeplabv3plus-xception 166.849G 54.709M 32.18 deeplabv3plus-mobile 52.875G 5.814M 79.62 hrnetv2_w18 32.948G 9.642M 19.02 hrnetv2_w32 80.177G 29.547M 17.02 hrnetv2_w48 165.336G 65.861M 17.21 pspnet-resnet50 118.447G 46.716M 54.19 pspnet-mobilenet 6.034G 2.377M 121.64 Our Model 46.541G 27.370M 33.66

[0092] (5) Summary: To address the technical challenges of fine-grained segmentation of icing on transmission lines under complex backgrounds, this paper proposes an improved semantic segmentation framework. Based on the lightweight and efficient SegFormer architecture, this framework introduces a multi-scale feature enhancement U-shaped network (MSFE-USN) and integrates three key components: a U-shaped network (USN) for deep fusion of encoder hierarchical features, an edge-guided multi-scale attention module (EMSAM) to enhance the perception and segmentation of subtle edges of icing on conductors, and a convolutional attention module (CBAM) to enhance local features and suppress irrelevant noise in complex backgrounds. Ablation experiments and comparative verification were conducted on a newly constructed private dataset. The results show that each module makes a significant positive contribution to the segmentation performance, and their combined use has a significant synergistic enhancement effect. The final model outperforms the original SegFormer and other mainstream comparative models in the three core metrics: mean intersection-union ratio (mIoU=94.07%), mean pixel accuracy (mPA=97.37%), and mean precision (mPrecision=96.4%). In extremely complex scenarios such as snow-covered valleys, lake reflections, nighttime rain and snow, and dense fog, the proposed model demonstrates superior anti-interference performance, effectively reducing missegmentation and missed segmentation, and improving stability and reliability in practical applications. In the task of fine-grained segmentation of iced areas and backgrounds, it can achieve continuous, complete, and high-precision extraction of icing contours, accurately distinguishing iced areas from complex backgrounds, and fully meeting the high standards for segmentation detail required by engineering applications. This research provides a robust and scalable solution for fine-grained semantic segmentation of transmission line icing, and also provides a useful reference for high-precision segmentation tasks in similar complex scenarios. Future work will focus on exploring model lightweighting and inference acceleration technologies to promote the real-time application of the algorithm in practical scenarios such as online monitoring of transmission line icing and deployment on mobile devices.

Claims

1. A method for refined segmentation of icing on transmission lines in the face of complex background interference, characterized in that, Includes the following steps: S1. Construct a semantic segmentation dataset for icing of transmission lines. Clean the original transmission line images, perform pixel-level annotation and dataset partitioning to form a basic dataset for model training and testing, and construct a generalization verification dataset across regions and devices. S2. Based on the three-segment network structure of encoder-feature enhancement module-decoder, a semantic segmentation model for icing of transmission lines is built. The encoder and decoder are based on the SegFormer semantic segmentation model architecture and are used to extract and reconstruct multi-scale semantic features. The feature enhancement module is a multi-scale feature enhancement U-shaped network MSFE-USN, which is used to suppress complex background interference and enhance fine-grained features of icing. S3. Train the transmission line icing semantic segmentation model using the training set in the basic dataset, and perform hyperparameter tuning and performance monitoring using the validation set in the basic dataset to obtain the trained transmission line icing semantic segmentation model. S4. Input the transmission line image to be segmented, which contains a complex background, into the trained transmission line icing semantic segmentation model, and output the pixel-level semantic segmentation result.

2. The method according to claim 1, characterized in that, The construction methods of the basic dataset mentioned in step S1 include: Original power transmission line images covering different time periods, weather conditions, and geographical backgrounds were collected, and the original power transmission line images were standardized and cleaned to remove invalid samples. The cleaned transmission line images were manually annotated at the pixel level, and three semantic categories were defined: background, main view transmission line, and side view transmission line. The labeled transmission line images are randomly divided into training subsets, validation subsets, and basic test subsets according to a preset ratio, and each subset is ensured to maintain a balanced distribution in terms of time, weather, background, and icing status.

3. The method according to claim 1, characterized in that, The construction method of the generalization validation dataset mentioned in step S1 includes: Collect images of power transmission lines that differ from the base dataset in terms of camera model, geographic background type, and icing physical type. The collected transmission line images with differences are cleaned and labeled at the pixel level, and the labeling specifications are consistent with the basic dataset. All labeled transmission line images are used as an independent generalization validation test set, which is not included in the model training process, and is used to evaluate the model's cross-domain generalization ability.

4. The method according to claim 1, characterized in that, The multi-scale feature enhancement U-shaped network MSFE-USN described in step S2 adopts a symmetrical U-shaped structure, which includes a top-down upsampling path and several lateral skip connection paths. The encoder outputs four layers of feature maps, denoted as F1, F2, F3, and F4, with their resolution decreasing sequentially. MSFE-USN processes these feature maps using the following process: S21. Deep feature processing: The deepest feature map F4 is input into a convolutional attention module CBAM for preliminary calibration, and then the calibrated feature map is input into an edge-guided multi-scale attention module EMSAM to output an enhanced feature map F5 that combines high-level semantics and rich edge details. S22. Feature fusion and upsampling: Upsample the enhanced feature map F5 by 2x bilinear interpolation to match the resolution of the feature map F3, and then fuse it with the shallow feature map F3 after CBAM enhancement. The number of channels is unified by 1×1 convolution on the fusion result, and the steps of upsampling, fusion with the shallow feature maps F2 and F1 enhanced by CBAM and channel unification are repeated. The calculation method for feature fusion is as follows: ; ;in, This represents the feature map of the i-th layer output by the encoder. This represents the intermediate enhanced feature map obtained during the progressive upsampling and fusion process of MSFE-USN, used to carry the fusion result of deep semantic information and corresponding shallow detail information. Up(·) indicates 2x bilinear interpolation upsampling. This represents a 1×1 convolution operator. This represents the convolutional attention module. This indicates an edge-guided multi-scale attention module; S23. Feature Output: Output a set of enhanced multi-scale feature maps and provide them to the decoder for final segmentation prediction.

5. The method according to claim 4, characterized in that, The Convolutional Attention Module (CBAM) processes the input feature map sequentially through the channel attention submodule and the spatial attention submodule; The channel attention submodule first performs global average pooling and global max pooling on the input feature map simultaneously to obtain two one-dimensional feature vectors; the two one-dimensional feature vectors are then fed into a shared two-layer multilayer perceptron for processing; the two processed one-dimensional feature vectors are added element by element, and then a channel attention weight vector is generated by passing the sigmoid activation function. The spatial attention submodule receives the feature map calibrated by the channel attention weight vector as input. First, it performs global average pooling and global max pooling along the channel dimension to obtain two two-dimensional feature maps. The two two-dimensional feature maps are then concatenated along the channel dimension. The concatenated two-dimensional feature map is then convolved through a 7×7 convolutional layer and then activated by the Sigmoid activation function to generate a spatial attention weight map.

6. The method according to claim 4, characterized in that, The edge-guided multi-scale attention module (EMSAM) is composed of the following cascaded units: Feature extraction and fusion unit: It contains a convolutional module for extracting input features; then, max pooling and average pooling operations are performed in parallel on the extracted features, and the two pooling results are concatenated along the channel dimension for preliminary fusion of local details and global context information; Feature enhancement and filtering unit: Connected after the feature extraction and fusion unit, the concatenated features are processed through a convolutional layer and then a feature weight map is generated by the Sigmoid activation function; The feature weight map is multiplied element-wise with the original convolution output of the feature extraction and fusion unit to perform preliminary feature enhancement and noise filtering. Multi-scale context aggregation unit: connected after the feature enhancement and filtering unit, it is a dilated spatial convolutional pooling pyramid (ASPP) structure; the ASPP structure contains at least three parallel branches, each using a 3×3 dilated convolutional layer with different dilation rates, to simultaneously capture global and local contextual information under different receptive fields, and to enhance edge feature responses at multiple scales. Two-dimensional fine calibration unit: connected after the multi-scale context aggregation unit, it is a parallel spatial and channel squeezing excitation scSE module; The scSE module runs a channel attention submodule cSE and a spatial attention submodule sSE in parallel, and adds the output feature maps of the two submodules after calibration along the channel direction to perform synchronous calibration of the spatial dimension and channel dimension of the feature map.

7. The method according to claim 7, characterized in that, The multi-scale context aggregation unit also includes an image pooling branch, which sequentially performs global average pooling, 1×1 convolution, and upsampling operations.

8. The method according to claim 7, characterized in that, The parallel spatial and channel squeezing excitation scSE module in the dual-dimensional fine calibration unit has a channel attention submodule cSE that performs global average pooling on the input feature map and then processes it through a two-layer multilayer perceptron and a sigmoid function to generate a channel attention weight vector. Its spatial attention submodule sSE generates a spatial attention weight map by performing a 1×1 convolution on the input feature map and then processing it with the Sigmoid function.

9. The method according to claim 4, characterized in that, The multi-scale feature enhancement U-shaped network MSFE-USN employs a collaborative enhancement strategy of cascading CBAM and EMSAM on the deepest feature map link in the lateral skip connection path. This includes: on the lateral connection path corresponding to the deepest feature map F4, the deepest feature map F4 is first preliminarily calibrated in terms of channel and spatial dimensions through the CBAM module to suppress invalid noise and highlight task-related regions; then the output of the CBAM module is used as the input of the edge-guided multi-scale attention module EMSAM, which synchronously executes multi-scale dilated convolution and parallel spatial channel attention to aggregate global and local context and enhance icing edge details, forming a cascaded processing flow of first filtering and denoising, and then edge enhancement.

10. The method according to claim 1, characterized in that, The decoder described in step S2 is used to perform channel unification, upsampling and fusion processing on the multi-scale features output by the feature enhancement module to generate pixel-level probability maps corresponding to each semantic category, and obtain the final pixel-level semantic segmentation result through normalization processing.