Icing detection method based on multi-source data feature interaction

Through the multi-source data feature interaction ice detection method, combined with image enhancement, DEM data and meteorological data, the problems of low recognition accuracy and poor real-time performance in the ice detection of transmission lines are solved, and efficient ice detection in complex environments is achieved.

CN120279425AActive Publication Date: 2025-07-08NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202510753442.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-08
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The prior art has problems in the detection of ice-covered transmission lines, such as low recognition accuracy, poor real-time performance, high computing resource consumption and insufficient adaptability to complex environments. It is especially difficult to achieve efficient and accurate ice-covered recognition under complex climate conditions.

Method used

The ice-covered detection method based on multi-source data feature interaction is adopted, and the ice-covered image is processed through dynamic routing image enhancement autoencoder. Combined with DEM data and meteorological data, the local and global features of ice-covered are extracted using the Terrafuse-BiNet network, and the meteorological timing features are extracted through the TRAM-Net module, and finally the ice-covered type and thickness level are output through the multi-task prediction module.

Benefits of technology

It significantly improves the accuracy and real-time performance of ice-covering detection, can complete ice-covering detection stably and efficiently in complex environments, improves adaptability to meteorological conditions and terrain changes, and reduces computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279425A_ABST
    Figure CN120279425A_ABST
Patent Text Reader

Abstract

The invention discloses an icing detection method based on multi-source data feature interaction, and belongs to the technical field of image processing, and the method comprises the steps: inputting an original icing image into a dynamic routing image enhancement auto-encoder, and generating a high-quality enhanced image; extracting multi-level features of the image, wherein the multi-level features comprise local icing features, icing global features and meteorological time sequence features; after uniform scale mapping and feature precoding are carried out on the icing local features, the icing global features and the meteorological time sequence features, interactive fusion is carried out through a bidirectional cross-modal attention mechanism and feature splicing operation to output icing fusion features; the icing fusion features are input into a multi-task prediction module, deep features are extracted through a residual fusion encoder, task mapping is completed, and corresponding icing types and icing thickness levels are output; according to the invention, through meteorological time sequence information and feature fusion, efficient icing detection is realized, different climate and topographic conditions can be adapted, and the operation safety and stability of the power transmission line are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to an icing detection method based on multi-source data feature interaction. Background Technique

[0002] The icing recognition of transmission lines is a key technology in the power system. Especially in cold regions, icing occurs frequently, increasing the weight of the line, intensifying conductor galloping, leading to fractures or equipment damage, and affecting the safety and stability of the power system. Therefore, it is crucial to identify the icing situation in a timely and accurate manner. Traditional manual inspections rely on on-site operations, which are intuitive and reliable. However, due to the extensive transmission lines and complex environments, there are problems such as low efficiency, missed detections, and delays. In recent years, automated monitoring technologies have gradually become a research hotspot, mainly using means such as unmanned aerial vehicles, remote sensing technologies, lidar, and infrared imaging, and realizing the automatic recognition and analysis of icing through image processing, machine learning, and deep learning algorithms. These technologies have made certain progress, but still face some challenges. On the one hand, complex climate conditions such as snow, fog, and backlight will affect image clarity and reduce the recognition accuracy. On the other hand, the high cost of sample annotation and scarce data limit the training of deep learning models. In addition, the engineering site has high requirements for computing resources and response speed, which also pose higher requirements for the lightweight and real-time performance of the model. How to efficiently and accurately identify the icing situation in massive data is still a difficult point in the development of technology. Therefore, improving the recognition accuracy, reducing missed and false detections, and enhancing the real-time performance are the key issues that need to be urgently solved in the current icing recognition technology.

[0003] The image-based icing detection method installs image acquisition devices on poles or other monitoring facilities, obtains and analyzes icing images to judge the icing state of transmission lines. Common methods include the detection method based on visual sensors and K-SVD denoising, which can effectively remove image noise and improve the recognition accuracy, but has a high computational complexity and a large dependence on lighting conditions, affecting the real-time performance. Another method is to use an unmanned aerial vehicle equipped with a binocular camera to obtain multi-view images and combine deep learning for feature matching to measure the icing thickness. However, unmanned aerial vehicles are greatly affected by bad weather and consume a high amount of computing resources, making it difficult to achieve all-weather real-time monitoring. In addition, the method based on Canny edge detection and Hough transform can extract the icing contour, but is sensitive to background interference and is easily affected by environmental lighting and occlusion.

[0004] In the task of detecting icing on transmission lines based on deep learning, using a dynamic routing image enhancement autoencoder to process the original icing images can effectively enhance the dark information, improve the contrast, and avoid problems such as detail loss, artifacts, or overprocessing. Moreover, meteorological time-series features are extracted from meteorological data, providing more discriminative meteorological semantic information for the icing detection task, and effectively improving the model's perception ability of icing states in complex environments. Most image-based icing detection methods often ignore the influence of meteorological factors and terrain features, while these factors play an important role in the formation and thickness distribution of icing. Meteorological conditions such as temperature, humidity, and wind speed in different regions, as well as the terrain where the transmission line is located, will significantly affect the accumulation speed and distribution characteristics of icing. Therefore, combining deep learning models with meteorological data and terrain data can more accurately predict icing conditions and optimize the detection results. The convolutional neural network of deep learning can not only effectively identify the icing features in the image, but also combine multi-source data to analyze the influence of environmental factors on icing, thereby improving the detection accuracy. Making full use of deep learning technology and integrating external environmental information will significantly improve the performance of detecting icing on transmission lines. Summary of the Invention

[0005] The present invention provides a method for detecting icing on transmission lines to address the deficiencies in the background technology. The method includes the joint recognition of icing types and icing thickness levels, uses meteorological time-series information and feature fusion to improve the detection accuracy, and optimizes the icing detection method for complex meteorological conditions, terrain changes, and low visibility environments, enabling it to stably and efficiently complete the detection of icing on transmission lines in different environments.

[0006] The present invention adopts the following technical solutions:

[0007] An icing detection method based on multi-source data feature interaction specifically includes the following steps:

[0008] S1. Collect actual original icing images of the transmission line through the monitoring cameras on the transmission line, organize the images into an image sequence in chronological order, and simultaneously obtain DEM data and meteorological data;

[0009] S2. Input the normalized icing image into a dynamic routing image enhancement autoencoder and output the enhanced image;

[0010] S3. Input the enhanced image and the DEM data into the Terrafuse-BiNet dual-branch feature extraction network in parallel. Among them, the deformable multi-scale detail feature extraction branch extracts edge textures, and the height difference-guided attention perception branch fuses the terrain for spatial modeling. Finally, local icing features and global icing features are obtained respectively;

[0011] S4. Input the meteorological data into the TRAM-Net meteorological feature extraction module, adopt the local meteorological variation rate to guide the attention bias, complete the feature encoding and output the meteorological time series features;

[0012] S5. After the local and global icing features and the meteorological time series features are subjected to unified scale mapping and feature pre-encoding, they are interactively fused through a bidirectional cross-modal attention mechanism and a feature splicing operation, and finally the multi-modal icing fusion features with aligned scales are output;

[0013] S6. Input the icing fusion features into the multi-task prediction module, where the residual fusion encoder extracts the deep features, and then the task-aware multi-head self-attention module extracts the type features and thickness level features, and finally inputs them into the lightweight classifier to obtain the corresponding icing type and icing thickness level.

[0014] As a further preferred solution of the icing detection method based on multi-source data feature interaction of the present invention, in step S2, the enhanced image is obtained, which specifically includes the following steps: S2.1. Input the original icing image at time into the noise type discriminator branch of the dynamic routing image enhancement autoencoder. After passing through the convolutional layer, global average pooling layer and a fully connected layer in the branch, the class label with the largest probability value is output by the softmax function , indicating that the image contains obvious Gaussian noise, indicating that the image has problems of blurring or low contrast, indicating that the image has no excessive noise and only needs lightweight enhancement, and its calculation process is expressed as: wherein, represents a convolutional block composed of a convolutional layer with a convolutional kernel , stride 2 and a ReLU function, GAP represents the global average pooling layer, and FC represents the fully connected layer;

[0015] S2.2. The original icing image at time is simultaneously input into the shared feature encoder branch of the dynamic routing image enhancement autoencoder. This branch extracts the general features such as the middle and high-level semantic structures and edge textures of the image, and gradually compresses the image size to reduce redundant information. The encoder is composed of three convolutional blocks in total, and each convolutional block is composed of convolution, batch normalization and ReLU function, and some layers are with max pooling operations, and finally outputs a multi-channel shared feature ;

[0016] S2.3. The branch controller adopts the DecSwitch routing strategy, and according to the output label of the noise type discriminator , activate one of the three decoder branches and route the shared features output by the encoder to this decoder branch. When it is input into the denoising branch Decoder1, when it is input into the blur enhancement branch Decoder2, when it is input into the lightweight enhancement branch Decoder3. The DecSwitch routing strategy is defined as follows:

[0017] ;

[0018] S2.4. When the noise category discriminator outputs , the shared features enter the denoising branch Decoder1 that can enhance clear textures, edges, and local structures. Its structure includes an UpConv upsampling convolution module and a ResBlock residual block. The residual block consists of a convolution and a ReLU activation function. After passing through a three-channel output layer, the sigmoid activation function compresses the pixels to the [0, 1] interval, and finally outputs the enhanced image at time . The relevant calculation formula is as follows:

[0019]

[0020]

[0021] where indicates that the enhanced image is output by the denoising branch, indicates the sigmoid activation function, indicates convolution. The UpConv upsampling convolution module includes a ConvTranspose2D transposed convolution and convolution, represents the input features, and BN represents batch normalization;

[0022] S2.5. When the noise category discriminator outputs , the shared features enter the blur enhancement branch Decoder2. After passing through multiple UpConv blocks, convolution blocks, as well as ReLU functions and CEB (Contrast Enhance Block) contrast enhancement blocks, finally, the enhanced image is output through the sigmoid function. The relevant calculation formula is as follows:

[0023]

[0024]

[0025] Among them, represents the input feature, and SE represents the channel attention mechanism for enhancing key channels;

[0026] S2.6. When the noise category discriminator outputs , the shared feature is input into the lightweight enhancement branch Decoder3, and only slight brightness and contrast adjustments are made. After upsampling through multiple UpConv modules, finally, the enhanced image is output through convolution and the sigmoid function.

[0027] As a further preferred solution of the ice coverage detection method based on multi-source data feature interaction of the present invention, in step S3, the local ice coverage feature and the global ice coverage feature are obtained, which specifically include the following steps:

[0028] S3.1. The enhanced image is input into the dual-branch feature extraction network Terrafuse-BiNet. In the deformable multi-scale detail feature extraction branch, it is first input into the basic feature extraction layer, which includes ordinary convolution, batch normalization operation, ReLU activation function, and CBAM convolutional attention module to obtain the primary feature . The specific feature extraction process can be expressed by the following formula:

[0029]

[0030] The primary feature is input into the deformable convolutional network including gradient guidance and regularization constraints. The deformable convolution allows dynamic sampling of the input feature map at each output position . The sampling position is jointly determined by the fixed sampling offset and the learnable offset . The calculation process of the output feature is expressed as:

[0031]

[0032] Among them, represents the feature response at the position in the output feature map, represents the input primary feature map, represents the convolution kernel weight of the th sampling point, is the number of sampling points, The weight coefficient guided by the input gradient satisfies , and the specific formula is as follows:

[0033]

[0034] Among them, represents the gradient magnitude at the position on the input feature map, is a hyperparameter for controlling local enhancement, used to control the response degree to local gradient changes, and introduces an L2 regularization term to constrain the offset amplitude. The specific constraint is:

[0035]

[0036] Among them, represents the offset regularization loss, represents the square of the L2 norm of, and finally outputs a structure-adaptive feature map ;

[0037] S3.2. Input into the atrous spatial pyramid pooling module to extract multi-scale context information. This module consists of three parallel atrous convolution branches with different atrous rates. The calculation process of the branches is expressed as:

[0038]

[0039] Among them, represents different atrous rates, represents the atrous convolution operation with an atrous rate of for , represents the local features extracted at a certain atrous rate. Concatenate the three groups of output features in the channel dimension to generate a detailed feature map , and the calculation process is expressed as:

[0040]

[0041] Among them, Concat represents the concatenation operation in the channel direction, are the local feature representations of the three parallel atrous convolution branches respectively. Finally, perform a residual connection to output the icing local features. The specific operation is as follows:

[0042]

[0043] S3.3. Input the enhanced image and the DEM data into the elevation-guided attention perception branch of the dual-branch feature extraction network Terrafuse-BiNet. First, perform concatenation in the channel dimension to generate a terrain joint feature map , which is expressed as:

[0044]

[0045] Among them, represents the DEM elevation data at time t. Subsequently, normalized coordinate encoding is introduced, and the normalized row-column coordinate tensor of each pixel point is concatenated to to form the spatial perception joint feature map :

[0046]

[0047]

[0048]

[0049] Among them, and respectively encode the horizontal and vertical position information. Normalize means linearly scaling the row and column indices to the interval [0, 1], and Meshgrid generates the corresponding coordinate tensor. , respectively represent the height and width. , respectively represent the horizontal and vertical coordinates; is input into a lightweight convolutional encoder. The encoder consists of two convolution blocks. Each convolution block contains a convolutional layer, batch normalization, and a ReLU activation function to extract the preliminary spatial terrain fusion feature . Subsequently, through convolution, is mapped to generate a query matrix , a key matrix , and a value matrix :

[0050]

[0051] The height difference spatial deviation function is introduced as a position enhancement term to calculate the attention score matrix. The calculation process is expressed as:

[0052]

[0053] Among them, and are respectively the query and key vectors at . represents the transpose operation. is the dimension of the key vector. and The normalized coordinate differences of pixel positions and , is the elevation difference of pixel position . Denote , and as a linear weighted combination:

[0054]

[0055] Wherein, , , are adjustable hyperparameters with a range of [0.1, 10]. The attention score matrix is normalized by the softmax function to obtain the attention weight . Subsequently, the features of the position to be attended are weighted and summed according to the weight to generate the attention aggregation result of the position :

[0056]

[0057] Use residual connection to process and to obtain the global icing feature :

[0058]

[0059] Wherein, ∈(0, 1) is an adjustable fusion coefficient used to adjust the proportional balance between and .

[0060] As a further preferred solution of the icing detection method based on multi-source data feature interaction of the present invention, in step S4, obtain meteorological time series features, which specifically include the following steps:

[0061] S4.1. Input the multi-dimensional meteorological data at each moment into the meteorological feature extraction module TRAM-Net. First, enter the variable weighting unit, which consists of a linear mapping layer, a non-linear activation layer, and a task-guided attention unit. After linear transformation by the vector weight matrix and the bias term, generate an intermediate representation through the tanh activation, and then calculate the similarity with the learnable vector and normalize it to obtain the variable importance weight of the current moment, which is used to adjust the dynamic variable, and its calculation process is as follows:

[0062]

[0063]

[0064] Among them, is a trainable mapping matrix, is a bias term, is a task attention vector, represents the transpose operation, represents the th meteorological feature vector at the time step, is the weighted meteorological feature at the moment, The value range of is (0, 1);

[0065] S4.2. Use the position vector to inject position information into the weighted meteorological feature to form an embedding vector with time and space perception capabilities, and input into the local temporal dynamic part, and use a one-dimensional convolutional sliding window with a kernel of 3 to extract the local dynamic features within the window. The calculation process is expressed as:

[0066]

[0067]

[0068] Among them, represents the th convolutional kernel, represents the local dynamic feature at the moment;

[0069] S4.3. Based on the weighted meteorological feature define a sliding window centered at the moment and with a length of and calculate the square of the L2 norm between at the moment and the average value of the features within the sliding window to obtain the local meteorological variability which is used to represent the deviation degree of the current meteorological state from other time points within the window. The variability is defined as follows:

[0070] ;

[0071] S4.4. Input the local dynamic feature into the mutation-enhanced attention module, and additionally construct , , ​, and introduce as an attention bias term into the attention score for offset, so that the module focuses on the key change moments, and finally complete weight normalization and weighted aggregation. The calculation process is as follows:

[0072]

[0073]

[0074]

[0075] Among them, , , represents the mapping matrix, represents the transpose operation, represents the empirical scaling factor in the self-attention mechanism, represents the mutation guidance coefficient, represents the th attention output after fusion at the time step, represents the total length of the meteorological time series;

[0076] S4.5. Perform softmax normalization on the local meteorological mutation rates at all moments to obtain the time weight coefficients , and perform weighted sum processing on to obtain the meteorological time series features with a unified length. The specific calculation process is as follows:

[0077]

[0078]

[0079] As a further preferred solution of the ice coating detection method based on multi-source data feature interaction of the present invention, in step S5, feature fusion is performed on the ice coating local features, ice coating global features, and meteorological time series features extracted in S2 and S3 to obtain ice coating fusion features, which specifically include the following steps:

[0080] S5.1. Input the ice coating global feature , the ice coating local feature and the meteorological time series feature into three groups of independent fully connected layers respectively for unified coding mapping. The linear mapping process of each feature path is:

[0081]

[0082] Among them, Represents a linear mapping operation, and outputs the encoded icing global features with a unified dimension , icing local features and meteorological time series features ;

[0083] S5.2. Encoded features and perform spatial feature interaction through a bidirectional cross-modal attention mechanism, and its calculation process is expressed as:

[0084]

[0085]

[0086]

[0087] Among them, represents the multi-head attention mechanism, and outputs the concatenated spatial fusion features , then and meteorological time series features also perform the attention mechanism twice with each other as inputs, realizing two-way information interaction and fusion, and obtaining the preliminary fusion features of a unified length after concatenation and linear mapping ;

[0088] S5.3. Input the preliminary fusion features into the lightweight Transformer encoder. After normalization, multi-head attention, feed-forward network, and residual connection, the final output is the fusion feature .

[0089] As a further preferred solution of the icing detection method based on multi-source data feature interaction of the present invention, in step S6, the icing type and icing thickness level are obtained, which specifically includes the following steps:

[0090] S6.1. Input the fusion feature into the multi-task prediction module for normalization. The features of each channel first extract the mean and standard deviation through global average pooling, and then perform standardization processing within the channel dimension. A channel offset term is introduced during normalization for scale adjustment, and then a feature tensor of a unified scale is output ;

[0091] S6.2. Take Input residual fusion encoder, which is composed of three residual blocks stacked in sequence. Each residual block contains two fully connected layers, a GELU activation function and an SE channel attention module. After the features are activated by GELU, they are sent to the SE channel attention module, where global average pooling, two-layer fully connected and sigmoid normalization are performed in sequence to generate channel weights, and the channel weights are multiplied by the original features, and then residual connection is performed to obtain the output of the first residual block. The output of the first residual block is used as the input of the second residual block. After the three residual blocks act in sequence, an enhanced feature tensor after multi-layer fusion is output ;

[0092] S6.3. Input into the multi-head self-attention layer, which is composed of two groups of attention heads, used to extract ice-covering type features and ice-covering thickness level features respectively. Each group of attention structures generates , , , respectively through linear transformation, and perform similarity calculation in the spatial dimension to obtain attention scores. Subsequently, weighted summation is performed on the value vectors, and a task bias matrix is introduced in each group of attention structures. After the bias matrix participates in the correction of the attention scores, the ice-covering type features and the ice-covering thickness level features are output respectively;

[0093] S6.4. Input the ice-covering type features and the ice-covering thickness level features into the lightweight prediction module respectively. This module consists of a global average pooling layer and two fully connected layers. Among them, the ice-covering type branch outputs the category prediction result , and the ice-covering thickness level branch outputs the level distribution . The calculation process is as follows:

[0094]

[0095]

[0096] Among them, GAP represents the global average pooling operation, , are the fully connected weight matrices of the corresponding branches of the task. Finally, the output result is determined by the maximum probability. In addition, the loss function is jointly trained in a task-weighted form:

[0097]

[0098] Among them, , represent the cross-entropy loss, , represent adjustable task weights, satisfying 。

[0099] As a further preferred solution of an ice coating detection method based on multi-source data feature interaction according to the present invention, it includes a line ice coating detection method, and the line ice coating detection method includes: a preprocessing module, a dual-branch feature extraction module, a meteorological feature extraction module, a feature fusion module, and a multi-task prediction module;

[0100] Preprocessing module: Organize the original transmission line ice coating images into an ice coating image sequence in chronological order, perform normalization processing, and then input them into an image enhancement autoencoder. Activate the corresponding decoding branch as needed to achieve denoising, blur enhancement, or lightweight enhancement, and obtain the enhanced image;

[0101] Dual-branch feature extraction module: Input the enhanced image into the deformable multi-scale detail feature extraction branch of the dual-branch feature extraction network Terrafuse_BiNet, extract structure-adaptive features, and fuse dilated convolutions to obtain multi-scale local information to form an ice coating local feature map. At the same time, splice the enhanced image with DEM data in the elevation difference-guided attention perception branch, and construct a spatial joint feature in combination with coordinate encoding. After attention calculation and aggregation, generate an ice coating global feature that fuses terrain information;

[0102] Meteorological feature extraction module: Input the multi-dimensional meteorological data at each moment into the meteorological feature extraction module TRAM-Net, extract key variables, and fuse time position information. Perceive short-term trends through local convolutions, and capture meteorological fluctuation features by combining the local meteorological variation rate of a sliding window. Finally, generate meteorological time series features through global enhancement and weighted aggregation;

[0103] Feature fusion module: After encoding the ice coating global feature, the ice coating local feature, and the meteorological time series feature respectively, successively pass through bidirectional cross-scale attention fusion, time series interaction fusion, and Transformer encoding, and finally obtain an ice coating fusion feature with a unified dimension;

[0104] Multi-task prediction module: Normalize the ice coating fusion feature and input it into a residual fusion encoder to extract enhanced features, and then generate ice coating type features and ice coating thickness grade features through a task-aware multi-head attention module, and send them into lightweight prediction modules respectively. Finally, output the prediction results of the ice coating type and the ice coating thickness grade.

[0105] Compared with the prior art, the present invention adopts the above technical solutions and has the following beneficial effects:

[0106] 1. The line icing detection method of the present invention adopts a multi-branch image enhancement strategy of noise discrimination and routing decoding. It automatically identifies the type of image noise through a dynamic routing image enhancement autoencoder, and activates the corresponding denoising, blurring enhancement, or lightweight enhancement branches, effectively enhancing the dark part information, improving the contrast, and dealing with various image quality degradation problems such as Gaussian noise, blurred defocus, and slight interference, enhancing the image structure details, being able to improve the image quality in complex environments, avoiding redundant calculations and maintaining feature integrity, and being significantly superior to traditional single-branch enhancement methods, providing a clear, stable, and high-quality image basis for subsequent icing feature extraction;

[0107] 2. The line icing detection method of the present invention constructs a double-branch Terrafuse-BiNet network structure, realizing the dual perception ability of complex terrain and local details. This structure shows stronger spatial adaptability in dealing with situations such as height differences, perspective offsets, etc., significantly improving the texture recognition ability of the icing area, effectively solving the problems of traditional methods being sensitive to terrain disturbances and unstable detection accuracy, and providing a more robust technical support for the comprehensive perception of mountain transmission lines;

[0108] 3. The line icing detection method of the present invention effectively improves the model's discrimination ability of key meteorological variables and its response ability to sudden climate changes by introducing the TRAM-Net meteorological feature extraction module. This method can accurately capture the influence signals of meteorological factors in the time series evolution process, enhancing the modeling depth of the detection system for complex climate dynamics, and significantly improving the prediction accuracy of the icing process and the accuracy of early warning. BRIEF DESCRIPTION OF THE DRAWINGS

[0109] Figure 1 is the flow chart of the line icing detection method of the present invention;

[0110] Figure 2 is the overall model structure diagram of the line icing detection method of the present invention;

[0111] Figure 3 is the module structure diagram of the image enhancement autoencoder in the method of the present invention;

[0112] Figure 4 is the effect diagram of the enhanced image output by the image enhancement autoencoder in the method of the present invention;

[0113] Figure 5 is the influence diagram of meteorological time series features on the model performance in the line icing detection method of the present invention;

[0114] Figure 6 is the comparison schematic diagram of the influence of meteorological time series features on the classification results of the present invention; Figure 6 In (a) is the classification result diagram of the present invention without adding meteorological time series features; Figure 6Among them, (b) is the classification result diagram after adding meteorological time series features to the present invention;

[0115] Figure 7 is the accuracy comparison curve diagram of the comparative experiment of the Terrafuse - BiNet model of the present invention; Specific embodiments

[0116] In order to make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the application will be further elaborated in detail below with reference to the accompanying drawings. The described embodiments are only a part of the embodiments involved in the present invention. All non - innovative embodiments of other researchers in this field belong to the protection scope of the present invention. At the same time, for the step numbers in the embodiments of the present invention, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0117] In an embodiment of the present invention, a method for detecting line icing, as Figure 1 shown, includes the following steps:

[0118] S1. Collect the actual original icing images of the transmission line through the monitoring cameras on the transmission line, organize the images into an image sequence in chronological order, and simultaneously obtain DEM data and meteorological data;

[0119] S2. Input the normalized icing images into the dynamic routing image enhancement auto - encoder to output enhanced images;

[0120] S2.1. Input the original icing image at time into the noise type discriminator branch of the dynamic routing image enhancement auto - encoder. After passing through the convolutional layer, global average pooling layer and a fully - connected layer in the branch, the category label with the largest probability value is output by the softmax function , , indicating that the image contains obvious Gaussian noise, indicating that the image has problems of blurring or low contrast, indicating that the image has no excessive noise and only needs light enhancement. Its calculation process is expressed as:

[0121]

[0122] wherein, represents a convolutional block composed of a convolutional layer with a convolutional kernel and a step size of 2 and the ReLU function, GAP represents the global average pooling layer, and FC represents the fully - connected layer;

[0123] S2.2. The original icing image at At the same time, it is input into the shared feature encoder branch of the dynamic routing image enhancement autoencoder. This branch extracts general features such as the middle and high-level semantic structures, edge textures of the image, and gradually compresses the image size to reduce redundant information. The overall encoder consists of three convolutional blocks, and each convolutional block is composed of convolution, batch normalization, and the ReLU function. Some layers have max pooling operations, and finally, a multi-channel shared feature ;

[0124] S2.3. The branch controller adopts the DecSwitch routing strategy. According to the output label of the noise type discriminator , one of the three decoder branches is activated, and the shared feature output by the encoder is routed to this decoder branch. When , it is input into the denoising branch Decoder1. When , it is input into the blur enhancement branch Decoder2. When , it is input into the lightweight enhancement branch Decoder3. The DecSwitch routing strategy is defined as follows:

[0125]

[0126] S2.4. When the noise category discriminator outputs , the shared feature enters the denoising branch Decoder1 that can enhance clear textures, edges, and local structures. Its structure includes the UpConv upsampling convolution module and the ResBlock residual block. The residual block consists of convolution and the ReLU activation function. After passing through a three-channel output layer, the sigmoid activation function compresses the pixels to the interval [0, 1], and finally, the image enhanced at time is output. The relevant calculation formula is as follows:

[0127]

[0128]

[0129] Among them, represents that this enhanced image is output by the denoising branch, represents the sigmoid activation function, represents convolution. The UpConv upsampling convolution module includes ConvTranspose2D transposed convolution and convolution. The input feature is denoted as, and BN represents batch normalization;

[0130] S2.5. When the noise class discriminator outputs the shared feature enters the fuzzy enhancement branch Decoder2. After passing through multiple UpConv blocks, convolutional blocks, the ReLU function, and the CEB (Contrast Enhance Block) contrast enhancement block, the enhanced image is finally output through the sigmoid function , and the relevant calculation formula is as follows:

[0131]

[0132]

[0133] Among them, denotes the input feature, and SE represents the channel attention mechanism for enhancing key channels;

[0134] S2.6. When the noise class discriminator outputs the shared feature is input into the lightweight enhancement branch Decoder3 for only slight brightness and contrast adjustment, and is upsampled through multiple UpConv modules. Finally, the enhanced image is output through convolution and the sigmoid function.

[0135] S3. The enhanced image and the DEM data are input into the dual-branch feature extraction network Terrafuse-BiNet in parallel. Among them, the deformable multi-scale detail feature extraction branch performs edge texture extraction, and the height difference-guided attention perception branch fuses the terrain for spatial modeling. Finally, the local icing features and the global icing features are obtained respectively, which specifically include the following steps:

[0136] S3.1. The enhanced image is input into the dual-branch feature extraction network Terrafuse-BiNet. In the deformable multi-scale detail feature extraction branch, it is first input into the basic feature extraction layer, which includes ordinary convolution, batch normalization operation, the ReLU activation function, and the CBAM convolutional attention module to obtain the primary feature . The specific feature extraction process can be expressed by the following formula:

[0137]

[0138] The primary feature is input into the deformable convolutional network including gradient guidance and regularization constraints. The deformable convolution allows at each output position Dynamically sample the input feature map, where the sampling positions are determined by a fixed sampling offset and a learnable offset jointly. The calculation process of the output features is expressed as:

[0139]

[0140] where represents the feature response at position in the output feature map, represents the input primary feature map, represents the convolutional kernel weight of the th sampling point, is the number of sampling points, is the weight coefficient guided by the input gradient, satisfying , and the specific formula is as follows:

[0141]

[0142] where represents the gradient magnitude at position on the input feature map, is a hyperparameter for controlling local enhancement, used to control the response degree to local gradient changes, and an L2 regularization term is introduced to constrain the offset magnitude. The specific constraint is:

[0143]

[0144] where represents the offset regularization loss, represents the squared L2 norm of, and finally output a structure-adaptive feature map ;

[0145] S3.2. Input into an atrous spatial pyramid pooling module to extract multi-scale context information. This module consists of three parallel atrous convolution branches with different atrous rates. The calculation process of the branches is expressed as:

[0146]

[0147] where represents different atrous rates, represents the atrous convolution operation with an atrous rate of for , represents the local features extracted at a certain atrous rate. Concatenate the three groups of output features in the channel dimension to generate a detailed feature map , and the calculation process is expressed as:

[0148]

[0149] Among them, Concat represents the splicing operation in the channel direction, which are the local feature representations of three parallel dilated convolution branches respectively, and finally a residual connection is performed to output the local icing features. The specific operations are as follows:

[0150]

[0151] S3.3. Input the enhanced image and the DEM data into the elevation-guided attention perception branch of the Terrafuse-BiNet dual-branch feature extraction network. First, perform channel dimension splicing to generate a terrain joint feature map , and its expression is:

[0152]

[0153] Among them, represents the DEM elevation data at time t. Subsequently, normalized coordinate encoding is introduced, and the normalized row and column coordinate tensors of each pixel point are spliced into to form a spatially aware joint feature map :

[0154]

[0155]

[0156]

[0157] Among them, and represent the encoded horizontal and vertical position information respectively. Normalize means linearly scaling the row and column indices to the interval [0, 1], and Meshgrid generates the corresponding coordinate tensors. and represent the height and width respectively. and represent the horizontal and vertical coordinates respectively. Input into a lightweight convolutional encoder. The encoder consists of two convolutional blocks. Each convolutional block contains a convolutional layer, batch normalization, and ReLU activation function to extract the preliminary spatially terrain fused features . Subsequently, through convolution, is mapped to generate a query matrix , a key matrix , and a value matrix :

[0158]

[0159] Introduce the height difference spatial deviation function As a position enhancement term, calculate the attention score matrix, and the calculation process is expressed as:

[0160]

[0161] Among them, and are respectively The query and key vectors at, Indicates the transpose operation, Is the key vector dimension, And Are the normalized coordinate differences of the pixel positions And respectively, Is the elevation difference of the pixel position , Indicates , And The linear weighted combination of:

[0162]

[0163] Among them, , , Are adjustable hyperparameters, with a range of [0.1, 10]. The attention score matrix After being normalized by the softmax function, the attention weight is obtained. Subsequently, the features of the position to be attended to are weighted and summed according to the weight to generate the attention aggregation result of the position :

[0164]

[0165] Use residual connection to process and to obtain the global icing feature :

[0166]

[0167] Among them, ∈(0, 1) is an adjustable fusion coefficient, used to adjust the and ratio balance between.

[0168] S4. Input the meteorological data into the TRAM-Net meteorological feature extraction module, adopt the local meteorological variation rate to guide the attention bias, complete the feature encoding and output the meteorological time series features, which specifically include the following steps:

[0169] S4.1. Input the multi-dimensional meteorological data at each moment into the TRAM-Net meteorological feature extraction module. First, enter the variable weighting unit, which consists of a linear mapping layer, a non-linear activation layer, and a task-guided attention unit. After linear transformation through the vector weight matrix and the bias term, generate an intermediate representation through tanh activation, and then calculate the similarity with the learnable vector and normalize it to obtain the variable importance weight at the current moment , which is used to adjust the dynamic variable, and its calculation process is as follows:

[0170]

[0171]

[0172] Among them, is the trainable mapping matrix, is the bias term, is the task attention vector, represents the transpose operation, represents the meteorological feature vector at the th time step, is the weighted meteorological feature at the moment, and the value range of

[0173] S4.2. Use the position vector to inject position information into the weighted meteorological feature to form an embedding vector with time and space perception ability. Input into the local time series dynamic part, and use a one-dimensional convolutional sliding window with a kernel of 3 to extract the local dynamic features within the window. The calculation process is expressed as:

[0174]

[0175]

[0176] Among them, represents the th convolutional kernel, represents the local dynamic feature at the

[0177] S4.3. Based on the weighted meteorological feature define a center at the moment and a length of The sliding window and calculate at the moment of the L2 norm squared between the average value of the features within the sliding window to obtain the local meteorological variation rate , which is used to represent the deviation degree of the current meteorological state from other time points within the window. The variation rate is defined as follows:

[0178] ;

[0179] S4.4. Input the local dynamic features into the mutation-enhanced attention module. Additionally, construct , , , and introduce as the attention bias term into the attention score for offsetting, so that the module focuses on the key change moments. Finally, complete the weight normalization and weighted aggregation. The calculation process is as follows:

[0180]

[0181]

[0182]

[0183] where , , represents the mapping matrix, represents the transpose operation, represents the empirical scaling factor in the self-attention mechanism, represents the mutation guidance coefficient, represents the th attention output after fusion at the time step, represents the total length of the meteorological time series;

[0184] S4.5. Perform softmax normalization on the local meteorological variation rates at all moments to obtain the time weight coefficients , and perform weighted sum processing on to obtain the meteorological time series features with a unified length. The specific calculation process is as follows:

[0185]

[0186]

[0187] After the local and global icing features and meteorological time-series features are mapped and pre-coded at a unified scale, they are interactively fused through a bidirectional cross-modal attention mechanism and feature splicing operation, and finally multi-modal icing fusion features with aligned scales are output, which specifically includes the following steps:

[0188] S5.1. Input the global icing feature , the local icing feature and the meteorological time-series feature into three groups of independent fully connected layers respectively for unified coding mapping. The linear mapping process of each feature path is as follows:

[0189]

[0190] Among them, represents the linear mapping operation, and the encoded global icing feature , the local icing feature and the meteorological time-series feature with unified dimensions are output;

[0191] S5.2. The encoded features and perform spatial feature interaction through a bidirectional cross-modal attention mechanism, and its calculation process is expressed as:

[0192]

[0193]

[0194]

[0195] Among them, represents the multi-head attention mechanism, and the spliced spatial fusion feature is output. Subsequently, and the meteorological time-series feature also serve as inputs to execute the attention mechanism twice to achieve bidirectional information interaction and fusion, and after splicing and linear mapping, the preliminary fusion feature with a unified length is obtained;

[0196] S5.3. Input the preliminary fusion feature into a lightweight Transformer encoder. After normalization, multi-head attention, feed-forward network, and residual connection, the fusion feature is finally output.

[0197] S6. Input the ice-covering fusion features into the multi-task prediction module. The residual fusion encoder extracts deep features, and then the task-aware multi-head self-attention module extracts the type features and thickness level features, and finally inputs them into the lightweight classifier to obtain the corresponding ice-covering type and ice-covering thickness level, which specifically includes the following steps:

[0198] S6.1. Input the fusion features into the multi-task prediction module for normalization. The features of each channel first extract the mean and standard deviation through global average pooling, and then perform normalization processing within the channel dimension. A channel offset term is introduced during the normalization process for scale adjustment, and then a feature tensor with a unified scale is output ;

[0199] S6.2. Input into the residual fusion encoder, which is composed of three residual blocks stacked in sequence. Each residual block contains two fully connected layers, a GELU activation function, and an SE channel attention module. After the features are activated by GELU, they are sent to the SE channel attention module, which sequentially performs global average pooling, two-layer fully connected, and sigmoid normalization to generate channel weights, multiplies the channel weights by the original features, and then performs residual connection to obtain the output of the first residual block. The output of the first residual block is used as the input of the second residual block. After the three residual blocks act in sequence, an enhanced feature tensor after multi-layer fusion is output ;

[0200] S6.3. Input into the multi-head self-attention layer, which is composed of two groups of attention heads, respectively used to extract the ice-covering type features and ice-covering thickness level features. Each group of attention structures respectively generates , , through linear transformation, and performs similarity calculation in the spatial dimension to obtain attention scores. Subsequently, weighted summation is performed on the value vectors. A task bias matrix is introduced in each group of attention structures, and after the bias matrix participates in the correction of the attention scores, the ice-covering type features and the ice-covering thickness level features are output respectively;

[0201] S6.4. Input the ice-covering type features and the ice-covering thickness level features into the lightweight prediction module respectively. This module consists of a global average pooling layer and two fully connected layers. Among them, the ice-covering type branch outputs the category prediction result , and the ice-covering thickness level branch outputs the level distribution . The calculation process is as follows:

[0202]

[0203]

[0204] Among them, GAP represents the global average pooling operation, 、 is the fully connected weight matrix of the corresponding branch of the task, and the output result is finally determined by the maximum probability. In addition, the loss function adopts a task-weighted form for joint training:

[0205]

[0206] Among them, 、 represents the cross-entropy loss, 、 represents the adjustable task weight, satisfying .

[0207] The overall structure of the present invention is as shown in Figure 2 and includes five core modules: a preprocessing module, a dual-branch feature extraction module, a meteorological feature extraction module, a feature fusion module, and a multi-task prediction module. The collaborative process is as follows:

[0208] First, the original icing image is obtained by the monitoring device and input into the preprocessing module to be organized into an icing image sequence in chronological order and subjected to normalization processing. Subsequently, it is input into the image enhancement autoencoder, the structure of which is as shown in Figure 3 , and the corresponding decoding branch is activated as needed according to the output result of the noise type discriminator to achieve denoising, blur enhancement, or lightweight enhancement, and an enhanced image is obtained. Figure 4 is the comparison effect diagram of the enhanced image and the original icing image.

[0209] The enhanced image is input into the deformable multi-scale detail feature extraction branch of the Terrafuse_BiNet dual-branch feature extraction network to extract structure-adaptive features and fuse dilated convolutions to obtain multi-scale local information to form an icing local feature map. At the same time, in the elevation difference-guided attention perception branch, the enhanced image is spliced with the DEM data and combined with coordinate encoding to construct a spatial joint feature. After attention calculation and aggregation, an icing global feature integrating terrain information is generated;

[0210] Meanwhile, the multi-dimensional meteorological data at each moment is input into the TRAM-Net meteorological feature extraction module to extract key variables and fuse time position information. The local convolution perceives short-term trends, and the sliding window mutation rate is combined to capture meteorological fluctuation features. Finally, meteorological time series features are generated through global enhancement and weighted aggregation;

[0211] After encoding the global icing features, local icing features, and meteorological time-series features respectively, they are successively passed through bidirectional cross-scale attention fusion, time-series interaction fusion, and Transformer encoding, and finally icing fusion features of a unified dimension are obtained;

[0212] Finally, the icing fusion features are normalized and then input into the residual fusion encoder to extract enhanced features. Then, through the task-aware multi-head attention module, branch features of the icing type and the icing thickness level are generated and respectively sent into the lightweight prediction module for final output of the prediction results of the icing type and the icing thickness level.

[0213] In addition, ablation experiments are conducted on the TRAM-Net meteorological feature extraction module. Experiments are respectively carried out in the models with and without adding meteorological time-series features and recorded, as shown in (a) in Figure 5 、 Figure 6 and (b) in Figure 6 as shown. The detailed results are shown in Table 1 below:

[0214] Table 1 Whether to add meteorological time series features Macro F1 Weighted F1 value Balanced accuracy Matthews correlation coefficient Accuracy No 88.1% 89.5% 88.3% 85% 89% Yes 89.5% 90.6% 89.7% 85.5% 90.6% It can be seen that after adding meteorological time-series features, the accuracy of the model is significantly improved, especially the predictions of glaze and rime are more accurate.

[0215] Meanwhile, advanced models such as MedViT, ResNet-101, EDPNet, and YOLO12 are used to train and test on the icing type recognition dataset. The performance of these models and the accuracy change curves on the validation set are recorded and compared with the method Terrafuse-BiNet of the present invention. The experimental results are as shown in Figure 7 shown. TerraFuse-BiNet reaches the optimal level in terms of Top-1 accuracy, precision, recall, and F1 comprehensive score, while maintaining a medium floating-point operations per second and good inference speed, demonstrating the structural design advantage and deployment balance in the complex icing image recognition task. The detailed results are shown in Table 2 below:

[0216] Table 2

[0217] Top-1 accuracy (%) Precision (%) Recall (%) F1 score (%) Floating point operations per second (G) Inference time (s) MedViT 88.20 87.41 86.47 86.99 1.1 0.0015 ResNet-101 87.60 87.00 85.92 86.44 0.65 0.0012 EDPNet 85.90 85.48 83.66 84.45 0.35 0.0007 YOLO12 83.89 84.52 82.90 83.71 0.2 0.0005 Terrafuse-BiNet (the method of the present invention) 89.01 88.31 86.97 87.1 0.8 0.001

[0218] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An icing detection method based on multi-source data feature interaction, characterized in that It includes the following steps: S1. Collect the actual original icing images of the transmission line through the monitoring cameras on the transmission line, organize the images into an image sequence in chronological order, and obtain DEM data and meteorological data at the same time; S2. Input the normalized icing image into the dynamic routing image enhancement autoencoder to output the enhanced image; S3. Input the enhanced image and the DEM data into the dual-branch feature extraction network Terrafuse-BiNet in parallel. The deformable multi-scale detail feature extraction branch extracts edge textures, and the elevation difference-guided attention perception branch fuses the terrain for spatial modeling to obtain the local icing features and the global icing features respectively; S4. Input the meteorological data into the meteorological feature extraction module TRAM-Net, use the local meteorological variation rate to guide the attention bias, complete feature encoding and output the meteorological time series features; S5. After the local icing features, global features and meteorological time series features are subjected to unified scale mapping and feature pre-encoding, they are interactively fused through the bidirectional cross-modal attention mechanism and feature splicing operation to output the multi-modal icing fusion features with aligned scales; S6. Input the icing fusion features into the multi-task prediction module. The residual fusion encoder extracts deep features, and then the task-aware multi-head self-attention module extracts the type features and thickness level features, and inputs them into the lightweight classifier to obtain the corresponding icing type and icing thickness level.

2. The icing detection method based on multi-source data feature interaction according to claim 1, characterized in that: In step S2, to obtain the enhanced image, it includes the following steps: S2.

1. Input the original icing image at the moment into the noise type discriminator branch of the dynamic routing image enhancement autoencoder. After passing through the convolutional layer, global average pooling layer, and a fully connected layer within the branch, the softmax function outputs the class label with the highest probability value . , indicating that the image contains obvious Gaussian noise, indicating that the image has blurring or low contrast problems, indicating that the image has no excessive noise and only requires lightweight enhancement, and its calculation process is expressed as: Among them, represents the convolution kernel , a convolution block composed of a convolution layer with a stride of 2 and a ReLU function, GAP represents the global average pooling layer, and FC represents the fully connected layer; S2.2、 Original icing image at a moment Meanwhile, it is input into the shared feature encoder branch of the dynamic routing image enhancement autoencoder. This branch extracts the general features of the image, including middle and high-level semantic structures and edge textures, and gradually compresses the image size to reduce redundant information. The overall encoder consists of three convolutional blocks, and each convolutional block is composed of convolution, batch normalization, and ReLU functions. Some layers have max pooling operations, and finally, a multi-channel shared feature is output ; S2.

3. The branch controller adopts the DecSwitch routing strategy and activates one of the three decoder branches according to the output label of the noise type discriminator, and routes the shared features output by the encoder to this decoder branch. When is input to the denoising branch Decoder1, when is input to the blurring enhancement branch Decoder2, when is input to the lightweight enhancement branch Decoder3. The DecSwitch routing strategy is defined as follows: when is input to the lightweight enhancement branch Decoder3. The DecSwitch routing strategy is defined as follows: ; S2.

4. When the noise category discriminator outputs , the shared feature enters the denoising branch Decoder1 that can enhance clear textures, edges, and local structures. Its structure includes an upsampling convolution module UpConv and a residual block ResBlock. The residual block consists of a convolution and a ReLU activation function. After passing through a three-channel output layer, the sigmoid activation function compresses the pixels to the interval [0, 1], and finally outputs the enhanced image at time . The relevant calculation formula is as follows: ; Among them, represents the enhanced image output by the denoising branch, represents the sigmoid activation function, represents convolution, the UpConv upsampling convolution module includes ConvTranspose2D transposed convolution and convolution, represents the input feature, and BN represents batch normalization; S2.

5. When the noise category discriminator outputs the shared feature enters the fuzzy enhancement branch Decoder2, passes through multiple UpConv blocks, convolutional blocks, ReLU functions, and CEB (Contrast Enhance Block) contrast enhancement blocks, and finally outputs the enhanced image through the sigmoid function. The relevant calculation formula is as follows: ; Among them, represents the input feature, and SE represents the channel attention mechanism for enhancing key channels; S2.

6. When the noise category discriminator outputs , the shared feature is input into the lightweight enhancement branch Decoder3, and only slight brightness and contrast adjustments are made. After upsampling through multiple UpConv modules, finally, the enhanced image is output through convolution and the sigmoid function.

3. The ice covering detection method based on multi-source data feature interaction according to claim 1, characterized in that: In step S3, to extract the local icing features and the global icing features, it includes the following steps: S3.

1. Input the enhanced image into the dual-branch feature extraction network Terrafuse-BiNet. In the deformable multi-scale detail feature extraction branch, it is first input into the basic feature extraction layer, which includes ordinary convolution, batch normalization operation, ReLU activation function, and CBAM convolutional attention module, to obtain primary features . The specific feature extraction process can be expressed by the following formula: Input the primary features into the deformable convolutional network with gradient guidance and regularization constraints. Deformable convolution allows dynamic sampling of the input feature map at each output position . The sampling position is jointly determined by the fixed sampling offset and the learnable offset . The calculation process of the output feature is expressed as: where, represents the feature response at position in the output feature map, represents the input primary feature map, represents the convolution kernel weight of the th sampling point, is the number of sampling points, is the weight coefficient guided by the input gradient, satisfying . The specific formula is as follows: where, represents the gradient magnitude at position on the input feature map, is a hyperparameter for controlling local enhancement, used to control the response degree to local gradient changes, and an L2 regularization term is introduced to constrain the offset amplitude. The specific constraint is: where, represents the offset regularization loss, represents the square of the L2 norm of, and finally output a structure-adaptive feature map ; S3.

2. Extract the multi-scale context information by inputting it into the Atrous Spatial Pyramid Pooling (ASPP) module. The ASPP module consists of three parallel atrous convolutional branches with different atrous rates. The calculation process of the branches is expressed as: Input it into the Atrous Spatial Pyramid Pooling (ASPP) module to extract multi-scale context information. The ASPP module is composed of three parallel atrous convolutional branches with different atrous rates. The calculation process of the branches is as follows: Among them, represents different atrous rates, represents the atrous convolution operation with an atrous rate of of the atrous convolution operation, represents the local features extracted at a certain atrous rate. Concatenate the three groups of output features in the channel dimension to generate the detailed feature map , and the calculation process is expressed as: Among them, Concat represents the concatenation operation in the channel direction, are the local feature representations of the three parallel atrous convolutional branches respectively. Finally, perform a residual connection to output the ice-covered local features. The specific operation is as follows: ; S3.

3. Input the enhanced image and the DEM data into the height difference-guided attention perception branch of the Terrafuse-BiNet dual-branch feature extraction network. First, perform channel dimension concatenation to generate a terrain joint feature map , which is expressed as: where represents the DEM elevation data at time t. Subsequently, introduce normalized coordinate encoding and concatenate the normalized row and column coordinate tensors of each pixel point to to form a spatially aware joint feature map : ; ; Among them, and respectively represent the encoded horizontal and vertical position information. Normalize means linearly scaling the row and column indices to the interval [0, 1], and Meshgrid generates the corresponding coordinate tensor. , respectively represent the height and width. , respectively represent the horizontal and vertical coordinates; is input into the lightweight convolutional encoder, and the encoder consists of two convolutional blocks. Each convolutional block contains a convolutional layer, batch normalization, and ReLU activation function to extract the preliminary spatial terrain fusion features . Subsequently, through convolution, is mapped to generate the query matrix , the key matrix , and the value matrix : Introduce the height difference spatial deviation function as the position enhancement term to calculate the attention score matrix. The calculation process is expressed as: Among them, and are respectively the query and key vectors at . represents the transpose operation. is the dimension of the key vector. and are respectively the normalized coordinate differences of the pixel positions and . is the elevation difference of the pixel position . represents , and 's linear weighted combination: Among them, , , are adjustable hyperparameters with a range of [0.1, 10]. The attention score matrix is normalized by the softmax function to obtain the attention weight . Subsequently, the features at the attended position are weighted and summed according to the weight to generate the attention aggregation result at the position : Use residual connection for and Processed to obtain the global icing characteristics : Among them, ∈(0,1) is an adjustable fusion coefficient used to adjust and The proportional balance between them.

4. The icing detection method based on multi-source data feature interaction according to claim 1, wherein In step S4, to extract the meteorological time series features, it includes the following steps: S4.

1. Input the multi-dimensional meteorological data at each moment into the TRAM-Net meteorological feature extraction module. First, enter the variable weighting unit, which consists of a linear mapping layer, a non-linear activation layer, and a task-guided attention unit. After linear transformation by the vector weight matrix and bias term, generate an intermediate representation through tanh activation. Subsequently, calculate the similarity with the learnable vector and normalize it to obtain the variable importance weight at the current moment , which is used to adjust the dynamic variable, and its calculation process is as follows: ; Among them, is the trainable mapping matrix, is the bias term, is the task attention vector, represents the transpose operation, represents the th meteorological feature vector at the time step, is the weighted meteorological feature at the moment, ranges from (0, 1); S4.

2. Use the position vector Inject the position information into the weighted meteorological features to form an embedding vector with time and space perception capabilities , and input into the local temporal dynamic part. Use a one-dimensional convolutional sliding window with a kernel of 3 to extract the local dynamic features within this window. The calculation process is expressed as: ; Among them, represents the th convolution kernel, represents the local dynamic feature at the S4.

3. Based on the weighted meteorological features Define a sliding window centered at the moment with a length of , and calculate the square of the L2 norm between the moment and the average value of the features within the sliding window to obtain the local meteorological variation rate , which is used to represent the deviation degree of the current meteorological state from other time points within the window. The variation rate is defined as follows: ; S4.

4. Input the local dynamic features into the mutation-enhanced attention module, and additionally construct 、 、 , and introduce as the attention bias term into the attention score for offset, so that the module focuses on the key change moments. Finally, complete weight normalization and weighted aggregation. The calculation process is as follows: ; ; Among them, , , represents the mapping matrix, represents the transpose operation, represents the empirical scaling factor in the self-attention mechanism, represents the mutation guidance coefficient, represents the attention output after fusion at the th time step, represents the total length of the meteorological time series; S4.

5. Softmax normalization is performed on the local meteorological variation rates at all times to obtain the time weight coefficients , and weighted sum processing is performed on to obtain the meteorological time series features of a unified length . The specific calculation process is as follows: ; .

5. A method for icing detection based on multi-source data feature interaction according to claim 1, characterized in that, In step S5, to obtain the icing fusion features, it includes the following steps: S5.

1. Input the global icing characteristics , the local icing characteristics and the meteorological time series characteristics into three independent fully connected layers respectively for unified coding mapping. The linear mapping process of each feature path is as follows: Among them, represents the linear mapping operation, and the encoded global icing characteristics , the local icing characteristics and the meteorological time series characteristics with a unified dimension are output; S5.2, Encoded Features and perform spatial feature interaction through a bidirectional cross-modal attention mechanism, and its calculation process is expressed as: ; ; Among them, represents the multi-head attention mechanism and outputs the spatially fused features after concatenation , and then and the meteorological time series features also perform the attention mechanism twice as each other's inputs to achieve two-way information interaction and fusion, and obtain the preliminary fused features of unified length after concatenation and linear mapping ; S5.

3. Input the preliminary fusion features into a lightweight Transformer encoder. After normalization, multi-head attention, feed-forward network, and residual connection, finally output the fusion features .

6. The icing detection method based on multi-source data feature interaction according to claim 1, wherein In step S6, to obtain the icing type and icing thickness level, it includes the following steps: S6.

1. Input the fusion feature into the multi-task prediction module for normalization. For each channel, the mean and standard deviation of the feature are first extracted through global average pooling, and then standardized within the channel dimension. A channel offset term is introduced during the normalization process for scale adjustment, and then a feature tensor of a unified scale is output ; S6.

2. Feed into the input residual fusion encoder, which is composed of three residual blocks stacked in sequence. Each residual block contains two fully connected layers, a GELU activation function, and an SE channel attention module. After being activated by GELU, the features are fed into the SE channel attention module, where global average pooling, two layers of fully connected layers, and sigmoid normalization are performed in sequence to generate channel weights. The channel weights are then multiplied by the original features, followed by residual connection to obtain the output of the first residual block. The output of the first residual block is used as the input of the second residual block. After the three residual blocks act in sequence, an enhanced feature tensor after multi-layer fusion is output. ; S6.

3. Input into the multi-head self-attention layer, which consists of two groups of attention heads for extracting icing type features and icing thickness level features respectively. Each group of attention structures generates , , and respectively through linear transformation, and performs similarity calculation in the spatial dimension to obtain attention scores. Subsequently, weighted summation is performed on the value vectors, and a task bias matrix is introduced in each group of attention structures. After the bias matrix participates in the correction of the attention scores, the icing type feature and the icing thickness level feature are output respectively;​​​​​​​​​​​​ S6.

4. Input the icing type feature and the icing thickness level feature into the lightweight prediction module respectively. This module consists of a global average pooling layer and two fully connected layers. The icing type branch outputs the category prediction result , and the icing thickness level branch outputs the level distribution . The calculation process is as follows: ; Among them, GAP represents the global average pooling operation, , is the fully connected weight matrix of the corresponding branch of the task, and the output result is finally determined by the maximum probability. In addition, the loss function adopts a task-weighted form for joint training: Among them, , represents the cross-entropy loss, , represents the adjustable task weight, satisfying .

7. A method for icing detection based on multi-source data feature interaction according to claim 1, characterized in that: It includes a preprocessing module, a dual-branch feature extraction module, a meteorological feature extraction module, a feature fusion module and a multi-task prediction module: Preprocessing module: Organize the original icing images of the transmission line into an icing image sequence in chronological order, perform normalization processing and then input them into the image enhancement autoencoder, activate the corresponding decoding branch as needed to achieve denoising, blur enhancement or lightweight enhancement, and obtain the enhanced image; Dual-branch feature extraction module: Input the enhanced image into the deformable multi-scale detail feature extraction branch of the dual-branch feature extraction network Terrafuse_BiNet, extract the structure adaptive features and fuse the dilated convolution to obtain multi-scale local information to form the local icing feature map. At the same time, in the elevation difference-guided attention perception branch, splice the enhanced image and the DEM data and combine the coordinate encoding to construct the spatial joint features. After attention calculation and aggregation, generate the global icing features that fuse the terrain information; Meteorological feature extraction module: Input the multi-dimensional meteorological data at each moment into the meteorological feature extraction module TRAM-Net, extract the key variables and fuse the time position information, perceive the short-term trend by local convolution, capture the meteorological fluctuation features by combining the sliding window local meteorological variation rate, and finally generate the meteorological time series features through global enhancement and weighted aggregation; Feature Fusion Module: After encoding the global icing features, local icing features, and meteorological time-series features respectively, they are sequentially passed through bidirectional cross-scale attention fusion, time-series interaction fusion, and Transformer encoding to finally obtain icing fusion features with a unified dimension; Multi-task Prediction Module: After normalizing the icing fusion features, they are input into a residual fusion encoder to extract enhanced features, and then task-aware multi-head attention modules are used to generate icing type features and icing thickness level features, which are respectively fed into lightweight prediction modules, and finally the prediction results of icing type and icing thickness level are output.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on semantic adaptive edge enhancement network

    CN118781596A

  • Line icing detection method based on deep learning

    CN119169535A

  • Remote sensing image change detection method based on adaptive Transform and deformable convolution

    CN119418204A

  • Person re-identification method and apparatus for fusing global features with ladder-shaped local features

    WO2024021394A1

Cited By

  • Power transmission line icing galloping prediction method and system based on improved time sequence large model, equipment and storage medium

    CN120724293A

  • Method and system for detecting icing thickness of overhead power line based on structure sensing framework

    CN120740462A

  • Multi-dimensional signal feature mapping method and device based on double-branch structure

    CN121188724A

  • Silicon carbide epitaxial growth real-time defect detection method

    CN121391847A

  • Arbitrary-scale super-resolution reconstruction method for multi-source heterogeneous DEM

    CN121437268A