Power line icing grade classification method, electronic equipment and storage medium

By combining the feature extraction methods of DConvLSTM, Swin Transformer and ResNet, the accuracy and real-time monitoring of ice-covered detection in complex environments are solved, and high-precision ice-covered area identification and thickness classification are realized, which is suitable for intelligent monitoring of power systems.

CN120495757APending Publication Date: 2025-08-15STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510583964.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing ice-covering detection technology has low accuracy and difficulty in real-time monitoring in large-scale and long-term monitoring. Especially in extreme weather, data collection is difficult and cannot meet the high frequency and high accuracy requirements of modern power systems for safety monitoring.

Method used

The DConvLSTM spatiotemporal feature extraction model is used to combine Swin Transformer and ResNet to extract the global and local features of the power line ice-covered image. Through spatial pyramid pooling fusion, multi-scale features are introduced, and terrain features are used to use enhanced multi-head self-attention and convolutional block attention module weighted enhancement features, combined with lightweight deconvolution module recovery resolution, and a high-precision ice-covered segmentation map is generated using a multi-loss collaborative optimization method, and a multi-modal fusion network is used to identify ice-covered type and thickness level classification.

Benefits of technology

It improves the accuracy and robustness of ice-covering detection, especially under complex terrain and variable meteorological conditions, and can more accurately identify ice-covering areas and thickness changes. It is suitable for large-scale data processing and real-time applications, providing high-quality ice-covering segmentation and thickness rating classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495757A_ABST
    Figure CN120495757A_ABST
Patent Text Reader

Abstract

The invention discloses a power line icing grade classification method, electronic equipment and a storage medium, belongs to the technical field of image processing, and solves the problem of how to improve the accuracy of power line icing detection.The method comprises the steps that firstly, an original icing image is obtained, and a standard icing image is preprocessed; a DConvLSTM spatio-temporal feature extraction model is used to extract meteorological spatio-temporal features, Swin Transform and ResNet are combined to extract global and local features of a standard icing image, and fusion of the global and local features of the image is realized through spatial pyramid pooling to obtain multi-scale features. The method comprises the following steps: performing weighted summation on meteorological spatial-temporal features, topographic features and multi-scale features to obtain a spatial-temporal multi-modal feature graph, performing block embedding by using a patch partition module, introducing a related mechanism and a CBAM module for weighted enhancement, and performing lightweight deconvolution, residual connection optimization and bilinear interpolation to obtain a high-quality graph. And the safety of a power line is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing and relates to a method for classifying ice coverage levels of power lines, electronic equipment and a storage medium. Background Art

[0002] Icing is a major challenge facing power systems in cold regions, especially in low-temperature, high-humidity weather conditions. Power lines, towers, and other electrical equipment are often affected by icing. Icing occurs when water vapor in the air rapidly condenses into a layer of ice upon encountering a surface below freezing. As the ice layer thickens, the mechanical load on the lines and equipment increases dramatically. In severe cases, this can lead to power line breakage, tower collapse, and equipment failure, seriously threatening the stable operation and reliable power supply of the power system. To effectively address this risk, the power industry urgently needs to monitor and assess the thickness of icing and its potential impact on electrical equipment in a timely and accurate manner.

[0003] Traditional icing detection technologies mainly include on-site measurements and sensor-based monitoring methods. For example, direct measurement and weighing methods directly measure the thickness or mass of the ice layer through manual or equipment to assess the degree of icing. Although these methods perform well in small-scale, high-precision detection, they have drawbacks such as large workload and difficulty in achieving large-scale automated monitoring. On the other hand, icing detection technologies based on conversion from other types of sensors, such as tension sensors and temperature sensors, have problems such as complex correspondences and low monitoring accuracy due to the conversion of other parameters into icing status. More importantly, traditional monitoring methods have difficulty achieving the goal of real-time icing monitoring under complex environmental conditions, especially in extreme weather conditions. Data collection is difficult and cannot meet the high-frequency and high-precision safety monitoring requirements of modern power systems.

[0004] Therefore, researchers have begun to explore new ice detection technologies based on remote sensing technology and image processing methods, in order to address the limitations of traditional methods in large-scale and long-term monitoring. In recent years, many studies have been devoted to developing automated detection methods based on image analysis. By collecting high-resolution images of power lines and equipment, combined with image processing and pattern recognition technology, the ice thickness classification can be achieved. Image-based ice detection methods have the advantages of high efficiency, low cost, and non-contact, and can improve the real-time monitoring of ice under complex environmental conditions. In addition, compared with traditional methods, the ice thickness classification method based on deep learning is more intelligent and has stronger adaptability. By continuously optimizing deep learning models and algorithms, deep learning-based ice detection technology will play an increasingly important role in the intelligent and automated monitoring of power systems in the future, providing more accurate and real-time technical support for the safe operation of power systems. Summary of the Invention

[0005] The technical solution of the present invention is used to solve the problem of how to improve the accuracy of ice coating detection on power lines.

[0006] The present invention solves the above technical problems through the following technical solutions:

[0007] The present invention provides a method for classifying ice coverage levels of power lines, comprising the following steps:

[0008] Performing preprocessing operations on the original ice-covered image to obtain a standard ice-covered image;

[0009] The DConvLSTM spatiotemporal feature extraction model is used to extract meteorological spatiotemporal features. The global and local features of the standard ice-covered image are extracted using Swin Transformer and ResNet, respectively. Spatial pyramid pooling is then used to fuse the global and local features of the image to obtain multi-scale features. Terrain features are then introduced, and the meteorological spatiotemporal features, terrain features, and multi-scale features are weighted and summed to obtain a spatiotemporal multimodal feature map.

[0010] The spatiotemporal multimodal feature map is input into the encoder composed of enhanced multi-head self-attention and deformable adaptive block, and the convolutional block attention module is combined to achieve feature weighted enhancement to obtain the enhanced encoded feature map;

[0011] A lightweight deconvolution module is used to restore image resolution. Residual connections, lightweight convolution, fused convolution, and an efficient channel attention mechanism are combined to optimize feature expression. Finally, bilinear interpolation is used to refine image features.

[0012] A multi-loss collaborative optimization method is used to generate high-precision ice segmentation maps for ice segmentation images.

[0013] The spatiotemporal multimodal feature map and ice area segmentation map are input into the multimodal fusion network to extract and optimize the segmentation feature map. The ice type is identified through the Softmax activation function, and the recognition result is optimized through the dynamic loss weight.

[0014] The ice segmentation results and ice type recognition results are input into the multi-task hybrid attention network, and the ice thickness level is preliminarily classified through the fully connected layer and attention mechanism. The dynamic weighted loss function is introduced to obtain the optimized ice thickness level classification, thereby obtaining the final ice thickness level result.

[0015] Furthermore, the method of performing preprocessing operations on the original ice-covered image to obtain a standard ice-covered image is as follows: obtaining the original ice-covered image, using an improved bilinear interpolation method and a Z-score detection method to process image missing values and outliers, performing wavelet denoising with soft and hard threshold fusion on the processed ice-covered image, using the CLAHE algorithm to perform enhanced adaptive histogram equalization processing, and obtaining a standard ice-covered image after normalization.

[0016] Furthermore, the calculation formula of the spatiotemporal multimodal feature map is as follows:

[0017]

[0018] in, represents the spatiotemporal multimodal feature map, F ERA5 represents the spatiotemporal characteristics of meteorology, F DEM Indicates terrain features, F SPP represents multi-scale features, and α, β, and γ all represent weight parameters.

[0019] Furthermore, the enhanced encoding feature map is represented as follows:

[0020] Z enhanced =Z c (i,j,d)+Z s (i,j,d)

[0021] Among them, Z enhanced Represents the enhanced encoding feature map, Z c (i, j, d) represents the result of channel attention optimization of the i-th row, j-th column, and d-th channel of the input image, Z s (i, j, d) represents the result of spatial attention optimization on the i-th row, j-th column, and d-th channel of the input image.

[0022] Furthermore, the method of restoring image resolution using a lightweight deconvolution module is as follows: using 1×1 convolution to compress the number of channels of the enhanced encoding feature map, reducing the number of channels to half of the original number, using 2×2 transposed convolution to enlarge the spatial size of the feature map to twice the original size, and then using another 1×1 convolution to restore the number of channels to the number of channels of the input feature map, thereby obtaining a spatial high-resolution feature map.

[0023] Furthermore, the total loss function of the generator in the multi-loss collaborative optimization method is expressed as follows:

[0024] L G =λ1·L adv +λ2·L Dice +λ3·L Focal

[0025] Among them, L G Denotes the total loss function of the generator, L adv represents the adversarial loss between the generator and the discriminator in the generative adversarial network framework, L Dice represents the Dice similarity loss between the generated image and the true label, L Focal Represents Focal loss, λ1, λ2, and λ3 represent the weight coefficients of the corresponding loss function.

[0026] Furthermore, the method of inputting the spatiotemporal multimodal feature map and the ice area segmentation map into the multimodal fusion network to extract and optimize the segmentation feature map, realizing ice type recognition through the Softmax activation function, and optimizing the recognition result through the dynamic loss weight is as follows:

[0027] The multimodal fusion network extracts and optimizes the initial ice segmentation map and performs classification output. The Softmax function is used to obtain the category probability distribution of each pixel.

[0028] The cross entropy loss function is used to calculate the difference between the prediction and the true label, and dynamic loss weights are added to balance the influence of different categories.

[0029] Furthermore, the ice segmentation results and ice type recognition results are input into the multi-task hybrid attention network, the ice thickness level is preliminarily classified through the fully connected layer and attention mechanism, and the dynamic weighted loss function is introduced to obtain the optimized ice thickness level classification, thereby obtaining the final ice thickness level result as follows:

[0030] The ice segmentation results and ice type recognition results are reduced in dimension through convolution operations, and the attention mechanism is applied in the channel dimension and spatial dimension respectively. The added together obtains the attention-enhanced fusion features, which are expressed as follows:

[0031] F hybrid =F c +F s

[0032] Among them, F hybrid represents the fusion feature after attention enhancement, F c represents the channel attention feature, F s Represents spatial attention characteristics;

[0033] The fusion feature F after attention enhancement hybrid Using a fully connected layer and a dynamic weighted loss function, the output is an optimized ice thickness level prediction result, which is expressed as follows:

[0034] Y=Softmax(FC(F hybrid ))

[0035] Among them, FC represents the fully connected layer and Y represents the probability distribution of each category.

[0036] The present invention also provides an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above-mentioned power line icing level classification method, and the processor is configured to execute the program stored in the memory.

[0037] The present invention also provides a storage medium having a computer program stored thereon. When the computer program is run by a processor, the steps of the above-mentioned method for classifying the ice coverage level of power lines are executed.

[0038] The beneficial effects of the present invention are as follows:

[0039] The method of the present invention uses the DConvLSTM spatiotemporal feature extraction model, combines meteorological data, global and local features of standard ice-covered images, and DEM (digital elevation model) data, and adopts spatial pyramid pooling (SPP) for fusion to obtain a spatiotemporal multimodal feature map. The spatiotemporal information of meteorological data, terrain characteristics, and spatial features of images are combined to enable the model to have a more comprehensive understanding of the ice-covered phenomenon. In particular, when considering the influence of terrain, the model can more accurately identify ice-covered areas and improve the accuracy of segmentation and thickness classification. In particular, it shows higher robustness when dealing with changes in ice thickness under complex terrain and changeable meteorological conditions.

[0040] The method of the present invention introduces enhanced multi-head self-attention (EMHSA) and convolutional block attention module (CBAM). The method adopts multiple attention mechanisms to weight and enhance the spatiotemporal multimodal feature maps. EMHSA automatically focuses on different key areas in the image through multiple attention heads; while CBAM optimizes features at the spatial and channel levels, so that the model can more accurately focus on the key feature areas of ice cover. This mechanism effectively solves the noise and interference information in ice cover images under complex meteorological conditions, so that the model can better distinguish different ice cover types and thicknesses, especially in the face of local details and large-scale background confusion. It can accurately identify the boundaries and thickness changes of ice cover, thereby significantly improving the accuracy of ice cover image segmentation.

[0041] The method of the present invention adopts a lightweight deconvolution module combined with a multi-feature optimization strategy to restore the resolution of the encoded feature map, which can accurately restore the high-resolution image while maintaining computational efficiency. The lightweight deconvolution module reduces the amount of calculation and parameters, improves computational efficiency, and is particularly suitable for large-scale data processing and real-time applications. At the same time, through the multi-feature optimization strategy, the model can extract key information from global and local features, accurately restore the details of the ice-covered area, and avoid information loss or blurring in traditional methods. This enables the model to provide more accurate input when generating ice segmentation images, providing high-quality feature maps for subsequent ice type identification and thickness grade classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a flow chart of the method for classifying ice coverage levels of power lines according to the present invention;

[0043] Figure 2 This is a diagram showing the overall model structure of the power line ice coverage grade classification method of the present invention;

[0044] Figure 3 This is a model structure diagram of Swin-EMHA enhanced feature extraction in line icing segmentation of the power line icing level classification method of the present invention;

[0045] Figure 4 This is a model structure diagram of lightweight deconvolution in line icing segmentation in the power line icing grade classification method of the present invention;

[0046] Figure 5 This is an effect diagram of the segmentation of the pre-processed ice image by the power line ice level classification method of the present invention;

[0047] Figure 6 This is a graph showing the loss and accuracy of the validation set of the ice thickness classification module of the power line ice classification method of the present invention;

[0048] Figure 7 Schematic diagram comparing the evaluation indicators of the power line icing level classification method of the present invention and other methods. DETAILED DESCRIPTION

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0050] The technical solution of the present invention is further described below with reference to the accompanying drawings and specific embodiments:

[0051] Example 1

[0052] like Figure 1 As shown, a method for classifying power line icing levels based on a cross-modal attention mechanism in an embodiment of the present invention includes the following steps:

[0053] 1. Preprocess the original ice-covered image to obtain a standard ice-covered image.

[0054] The preprocessing operation includes: an interpolation operation, an outlier detection operation, a noise reduction operation and an enhancement operation.

[0055] 1.1 Interpolation Operation

[0056] The improved bilinear interpolation method is used to process the original ice-covered image. The specific formula is as follows:

[0057]

[0058] Among them, α ij represents the weighting factor, f'(x,y) represents the pixel grayscale value of the target point after interpolation, ||d ij || represents the Euclidean distance between the interpolation point and its neighboring points, ∑ i,j exp(-||d ij || 2 / σ 2 ) represents the normalization of the weights of all neighboring points, σ represents the smoothing factor, f(i,j) represents the pixel grayscale value of the neighboring point, ω ij Represents α ij Normalized weight factor.

[0059] 1.2. Outlier Detection Operation

[0060] The local Z-score detection method is used to perform outlier detection on the interpolated image, thereby improving the accuracy of outlier detection in the original ice-covered image.

[0061] The formula for detecting outliers in the local Z-score detection method for the interpolated image is as follows:

[0062]

[0063] Among them, μ w represents the mean grayscale value of all data points in the neighborhood w where the pixel point is located, n represents the total number of data points in the neighborhood w, σ w represents the standard deviation of pixel values within the neighborhood w, Z w Represents the local Z-score value of pixel x, iw represents the pixel point in the neighborhood w, Indicates that the pixel point in the neighborhood w is i w Gray value, i w ∈w represents index i w Belongs to neighborhood w,i σ represents the pixel index within the neighborhood w, Indicates that the index in the neighborhood w is i σ Gray value, i σ ∈w represents index i σ Belongs to the neighborhood w, x w Represents the grayscale value of pixel x in the neighborhood w.

[0064] 1.3 Noise reduction operation

[0065] Perform wavelet decomposition on the image after the outlier detection operation to obtain the multi-scale wavelet decomposition coefficient. The calculation formula of the multi-scale wavelet decomposition coefficient is as follows:

[0066]

[0067] in, Represents the multi-scale wavelet decomposition coefficient, I(x,y) represents the image after the outlier detection operation, i scale Indicates the scale, j pos Indicates location, is a wavelet function.

[0068] For each wavelet decomposition coefficient w ij Perform soft and hard value threshold fusion to remove image noise. The specific formula is as follows:

[0069]

[0070] in, represents the wavelet decomposition coefficient after processing, λ represents the threshold, α represents the adjustment parameters of the hard threshold and soft threshold after fusion, sign(w ij ) represents the symbolic function, max(|w ij |-λ,0) represents the maximum value function.

[0071] The processed wavelet decomposition coefficients The denoised image is obtained through inverse wavelet decomposition transformation.

[0072] 1.4 Enhanced Operation

[0073] The CLAHE algorithm is used to enhance the denoised image, and each pixel (x, y) is processed by the local contrast weighted formula to obtain a standard ice-covered image.

[0074] The local contrast weighted formula is expressed as follows:

[0075]

[0076] Among them, I'(x,y) represents the enhanced pixel gray value, I CLAHE (x, y) represents the grayscale value of the pixel after processing by the CLAHE algorithm, I min Indicates the minimum grayscale value in the local window, I max Represents the maximum grayscale value in the local window, and W(x,y) represents the local contrast weighting coefficient.

[0077] 2. The DConvLSTM spatiotemporal feature extraction model is used to extract meteorological spatiotemporal features. The global and local features of the standard ice-covered image are extracted by combining Swin Transformer and ResNet, respectively. The global and local features of the image are fused through spatial pyramid pooling (SPP) to obtain multi-scale features. Then, terrain features are introduced, and the meteorological spatiotemporal features, terrain features, and multi-scale features are weighted and summed to obtain a spatiotemporal multimodal feature map.

[0078] 2.1. Input meteorological data into DConvLSTM spatiotemporal feature extraction model, and calculate weight γ from current meteorological data through MLP (Multi-Layer Perceptron) multi-layer perceptron t And in memory cell state C t Update in order to obtain the meteorological spatiotemporal characteristics F ERA5 , the updated memory unit state expression is as follows:

[0079]

[0080] Among them, f t Indicates the degree of forgetting of the previous time step by the current time step, C t-1 Represents the memory cell state at the previous time step, i t represents the input gate, Represents candidate memory information, and ⊙ represents element-by-element multiplication.

[0081] 2.2. Input the standard ice-covered image into the Swin Transformer and ResNet to extract the global features and local features respectively. After fusion, a mixed feature map F with a size of H×W×C is obtained. Pooling windows of different sizes are used on the mixed feature map F to obtain multi-scale features F. SPP , the formula is as follows:

[0082] F SPP =Conv(Concat(Pool 1×1 (F),Pool 2×2(F),Pool 4×4 (F)))

[0083] Among them, H, W and C are the height, width and number of channels of the mixed feature map F respectively. The Concat function represents the splicing of feature maps of different pooling windows. The Conv function represents the convolution operation, which is used to further fuse the spliced multi-scale features. k×k Indicates that a pooling operation is performed using a pooling window of k×k (k=1, 2, 4).

[0084] 2.3. Meteorological spatiotemporal characteristics F ERA5 、Topographic features F DEM And the multi-scale feature F SPP Perform weighted summation to obtain the spatiotemporal multimodal feature map, the formula is as follows:

[0085]

[0086] in, represents the spatiotemporal multimodal feature map, and α, β, and γ all represent weight parameters.

[0087] 3. The spatiotemporal multimodal feature map is input into the encoder composed of Enhanced Multi-Head Self-Attention (EMHSA) and Deformable Adaptive Patch Module (DeforAPM), and combined with the Convolutional Block Attention Module (CBAM) to achieve feature weighted enhancement and obtain the enhanced encoded feature map.

[0088] 3.1. Divide the spatiotemporal multimodal feature map processed in step 2 into non-overlapping P×P patches, flatten each two-dimensional patch into a one-dimensional vector, and combine all patches into a two-dimensional matrix X patch , through linear projection, the two-dimensional matrix X patch The features are embedded into a fixed-dimensional space with D=512 to obtain the embedded Patch feature matrix. The specific expression is as follows:

[0089] Z=X patch W e +b e

[0090] Among them, Z represents the embedded Patch feature matrix, W e represents the linear weight of the embedding, b e Represents the bias vector.

[0091] 3.2. Divide each patch in the feature matrix Z into H attention heads, and the dimension of each head is Use each head to calculate the query matrix Q, key matrix K and value matrix V respectively. For each head, use Q and K to calculate the relationship between each patch and other patches, and get the attention weight matrix as follows:

[0092]

[0093] Among them, A represents the attention weight matrix, Softmax is a normalization operation that converts the output into a probability distribution, and QK T Indicates the similarity between each patch and other patches, D h Indicates the dimensions of each head.

[0094] Use the attention weight matrix A to perform weighted summation on the value matrix V to obtain the enhanced result O of each head for the input feature, and add the outputs O1, O2, ..., O P×P Spliced together, and then passed through an output mapping matrix W o Processing, the enhanced feature matrix is obtained as follows:

[0095] Z'=Concat(O1,O2,...,O P×P )W o

[0096] Among them, Z' represents the enhanced feature matrix, and Concat represents the concatenation of P×P headers.

[0097] 3.3. Input the enhanced feature matrix Z' into the DeforAPM layer. Use the convolution operation to calculate the offset of each pixel point for each patch of the enhanced feature matrix Z'. Then dynamically change the sampling position of the patch according to the offset Δ to obtain the adjusted patch, which is expressed as follows:

[0098] P'=Deform(P,Δ)

[0099] Among them, P' represents the adjusted Patch, Deform represents the deformable convolution operation, and P represents the original Patch.

[0100] The offset Δ is dynamically expressed as follows:

[0101] Δx, Δy=Conv offset (Z')

[0102] Among them, Δx, Δy represent the pixel position, Conv offset Represents the convolutional layer that calculates the offset.

[0103] Through the pooling operation, P' is merged into multiple adjacent Patch features. The merged feature map is represented as follows:

[0104] Z merged =Pool(P',Stride=2)

[0105] Among them, Z merged Represents the merged low-resolution feature matrix, and Pool(P',Stride=2) represents a pooling operation with a stride of 2 performed on P'.

[0106] 3.4. Low-resolution feature matrix Z merged Perform global average pooling on each channel of , compress the spatial dimension, and only retain the global average value of each channel to obtain a vector of size 1×1×D. Then, the vector is processed by two layers of fully connected layers to obtain the attention weight of the output channel as follows:

[0107] M c =σ(W2·ReLU(W1·Pool avg (Z merged )))

[0108] Among them, M c Represents the attention weight of the output channel, W1 and W2 represent the weight matrices of the two fully connected layers respectively, σ is the Sigmoid activation function, and the ReLU activation function is used to increase the nonlinear expression ability of the network. Pool avg Represents a global average pooling operation.

[0109] The attention weight M of the output channel is obtained c With the low-resolution feature matrix Z merged Multiply each channel to strengthen the important channel features and obtain the channel-weighted feature Z c .

[0110] 3.5. Low-resolution feature matrix Z merged The channel dimension is compressed, and the low-resolution feature matrix Z is merged Perform average pooling and maximum pooling operations to obtain Z avg and Z max And spliced, the spliced feature map is input into a 7×7 convolution layer to extract the local spatial attention features, and the spatial attention weight is expressed as follows:

[0111] M s =σ(Conv 7×7 ([Z avg ,Z max ]))

[0112] Among them, M srepresents the spatial attention weight, [Z avg ,Z max ] represents the concatenated feature map, Conv 7×7 Represents a 7×7 convolutional layer operation.

[0113] The spatial attention weight M s With the low-resolution feature matrix Z merged Multiply, strengthen the important position, and obtain the spatially weighted feature Z s .

[0114] 3.6. Feature Z after channel weighting c and spatially weighted feature Z s By adding elements and fusing the attention information, the enhanced encoded feature map is expressed as follows:

[0115] Z enhanced =Z c (i,j,d)+Z s (i,j,d)

[0116] Among them, Z enhanced Represents the enhanced encoding feature map, Z c (i, j, d) represents the result of channel attention optimization of the i-th row, j-th column, and d-th channel of the input image, Z s (i, j, d) represents the result of spatial attention optimization on the i-th row, j-th column, and d-th channel of the input image.

[0117] 4. A lightweight deconvolution (UpSample) module is used to restore image resolution. Residual connections, 1×1 lightweight convolution, 3×3 fused convolution (LDFM), and efficient channel attention (ECA) mechanisms are combined to optimize feature expression. Finally, bilinear interpolation is used to refine image features.

[0118] 4.1. Use 1×1 convolution to compress the number of channels of the enhanced encoded feature map output in step 3, reducing the number of channels to half of the original number, as shown below:

[0119] F reduce =W reduce *F S3 +b reduce

[0120] Among them, F reduce Represents the compressed feature map, W reduce represents the 1×1 convolution weight matrix, F S3 represents the input feature map, b reduce represents the bias term.

[0121] Use 2×2 transposed convolution to double the spatial size of the feature map, and then use another 1×1 convolution to restore the number of channels to the input feature map F. S3 The number of channels is increased to obtain a spatial high-resolution feature map.

[0122] 4.2. Use residual connections to fuse the deep features of the encoder with the shallow features of the decoder to generate a residual fused feature map. Use deep convolution to extract spatial features channel by channel, then use point convolution to integrate inter-channel information at each pixel position, and finally add a bias to generate the output feature map, which is represented as follows:

[0123] F out =W point *(W depth *F residual )+b out

[0124] Among them, F out represents the output feature map, W point Represents the weight matrix of point convolution, the point convolution size is 1×1, W depth represents a 3×3 depth convolution kernel, F residual Represents the feature map after residual fusion, b out represents the bias term.

[0125] 4.3. Optimized feature map F along the spatial dimension out Perform global average pooling to compress the spatial dimension and retain only the global average value of each channel to obtain a one-dimensional vector F gap , dynamically determine the one-dimensional convolution kernel size K according to the number of channels C' of the encoded feature map, and use this one-dimensional convolution to F gap Extract the interaction weights between channels, expressed as follows:

[0126] F conv (c) = W 1D *F gap (c)+b conv

[0127] Among them, F conv (c) indicates that the Sigmoid activation function is used to convert F conv (c) Map to the interval [0, 1], generate channel attention weights, apply the attention weights of each channel to the original feature map of the corresponding channel, optimize the expression ability of each channel, and obtain the optimized feature map as follows:

[0128] F optimized (i, j, c) = A(c)·F out (i,j,c)

[0129] Among them, Foptimized (i, j, c) represents the value of the optimized feature map at the (i, j) spatial position of the c-th channel, A(c) represents the attention weight of the c-th channel, and F out (i, j, c) represents the value of the feature map after residual fusion at the (i, j) spatial position of the cth channel.

[0130] 4.4. Perform weighted summation on each channel of the optimized feature map point by point within the 3×3 local receptive field window, and the convolution feature map is represented as follows: The weight value of the c channel, W 1D represents the weight matrix of one-dimensional convolution, b conv Represents the bias term of one-dimensional convolution.

[0131]

[0132] Among them, F conv (i, j, c) represents the value of the optimized feature map at the (i, j) spatial position of the cth channel. represents the weight of the 3×3 convolution kernel of the c channel at its (m,n) position, b c is the bias term of the cth channel, F optimized (i+m,j+n,c) represents the value of the optimized feature map at the (i+m,j+n) spatial position of the cth channel.

[0133] The ReLU activation function is used to perform nonlinear mapping on the convolution feature map. The negative values after mapping are set to 0, and the positive values remain unchanged. The activated feature map F is obtained. ReLu .

[0134] 4.5. Use bilinear interpolation to convert the activated feature map F ReLu Upsampling to the target resolution is represented as follows:

[0135]

[0136] Among them, F up (x0, y0) represents the pixel value of the target position (x0, y0) after interpolation, represents the weight coefficient, Indicates the pixel value of the target position (x0, y0) rounded down.

[0137] 5. A multi-loss collaborative optimization method is used to generate a high-precision ice segmentation map for the ice segmentation image.

[0138] 5.1. Input the initial segmentation result map into the generator to generate the optimized segmentation map, which is expressed as follows:

[0139] F fake =G(F seg )

[0140] Among them, F seg represents the initial ice segmentation map, G represents the generator, and F fake Represents the optimized ice segmentation map.

[0141] The discriminator judges the input initial ice segmentation map F seg is the true ice segmentation label F real Or the optimized ice segmentation map F fake , which is expressed as follows:

[0142]

[0143] Where D(F) represents the output value of the discriminator. When the discriminator considers the input initial ice segmentation map to be the true ice segmentation label, the output value is 1. When the discriminator considers the input initial ice segmentation map to be the optimized ice segmentation map, the output value is 0.

[0144] 5.2. The optimization goal of the generator G is to minimize the adversarial loss, Dice loss, and Focal loss. The total loss function of the generator G is expressed as follows:

[0145] L G =λ1·L adv +λ2·L Dice +λ3·L Focal

[0146] Among them, L G Denotes the total loss function of the generator, L adv It represents the adversarial loss between the generator and the discriminator in the generative adversarial network framework, which is the loss generated by the generator and the discriminator during the game process; L Dice represents the Dice similarity loss between the generated image and the true label, which measures the overlap between the predicted segmentation result and the true label at the pixel level; L Focal Represents Focal loss, which handles the class imbalance problem through the Focal loss function; λ1, λ2, and λ3 represent the weight coefficients of the corresponding loss function respectively.

[0147] 6. Input the spatiotemporal multimodal feature map and the ice-covered area segmentation map into the multimodal fusion network to extract and optimize the segmentation feature map, realize ice cover type recognition through the Softmax activation function, and optimize the recognition result through the dynamic loss weight.

[0148] 6.1 Extracting and Optimizing the Initial Ice Cover Segmentation Map F Using Multimodal Fusion Network seg And perform classification output, use the Softmax function to get the category probability distribution of each pixel, which is expressed as follows:

[0149]

[0150] Among them, P(t ice1 |x) indicates that pixel x belongs to category t ice1 The probability of F seg (x,t ice1 ) indicates the pixel x in the feature map corresponds to category t ice1 Score, F seg (x,t ice ) indicates that pixel x in the feature map corresponds to other categories t ice The score, t ice Indicates that in addition to the current category t ice1 Other categories besides .

[0151] 6.2. Use the cross entropy loss function to calculate the difference between the prediction and the true label, and add dynamic loss weights to balance the influence of different categories, as shown below:

[0152]

[0153] in, Represents category t c Dynamic weight, Y true (i,t c ) represents the tth sample of the i-th c The label value of the category, P(t c |x i ) represents pixel x i Belongs to category t c The probability of loss, L represents the loss value, N represents the number of samples involved in calculating the loss, t sum Indicates the total number of categories.

[0154] 7. The ice segmentation results and ice type recognition results are input into the Multi-Task Hybrid Attention Network (MTHAnet). The ice thickness level is preliminarily classified through the fully connected layer and attention mechanism. A dynamic weighted loss function is introduced to obtain the optimized ice thickness level classification, thereby obtaining the final ice thickness level result.

[0155] 7.1. The ice segmentation results and ice type recognition results are reduced in dimension through convolution operations. Attention mechanisms are applied to the channel dimension and spatial dimension respectively. The attention-enhanced fusion features are added together and expressed as follows:

[0156] F hybrid =F c +F s

[0157] Among them, F hybridrepresents the fusion feature after attention enhancement, F c represents the channel attention feature, F s Represents spatial attention features.

[0158] 7.2. Fusion feature F after attention enhancement hybrid Using a fully connected layer and a dynamic weighted loss function, the output is an optimized ice thickness level prediction result, which is expressed as follows:

[0159] Y=Softmax(FC(F hybrid ))

[0160] Among them, FC represents the fully connected layer and Y represents the probability distribution of each category.

[0161] Example 2

[0162] The overall model structure of the present invention is as follows Figure 2 As shown, the model as a whole consists of four parts, and the specific construction process is as follows:

[0163] First, as described in step 1, the original ice-covered image is obtained and processed using the improved bilinear interpolation method and Z-score detection method. The processed image is de-noised using the wavelet denoising method with soft and hard threshold fusion. The image quality is improved by combining the enhanced contrast-limited adaptive histogram equalization (E-ClAHE) technique. The image is then normalized to obtain a standardized ice-covered image.

[0164] Then, according to step 2, the DConvLSTM module is used to extract the spatiotemporal dynamic features of the meteorological data. The Swin Transformer is used to extract the global features of the ice-covered image, and the ResNet is combined to extract the local detail features of the image. On this basis, the spatial pyramid pooling (SPP) module is used to fuse multi-scale features. The fused meteorological and terrain features are then spliced with the image features to generate a spatiotemporal multimodal feature map.

[0165] Then, according to step 3, the ice-covered image is divided into blocks to extract local features, such as Figure 3 As shown in Figure 2, the enhanced multi-head self-attention (EMHSA) mechanism is used to capture global dependencies, and the deformable adaptive block merging (VDSA) module is used to dynamically adjust the feature extraction strategy according to the regional complexity. The CBAM module is combined to perform weighted processing on important features. According to step 4, in the decoder stage, as shown in Figure 2, Figure 4As shown in Figure 2, a lightweight deconvolution (UpSample) module is used to restore the image resolution. The residual connection, 1×1 lightweight convolution, 3×3 fused convolution (LDFM) and ECA-Net attention mechanism are combined to optimize the feature expression. Finally, the image features are refined by bilinear interpolation. Then, according to step 5, a multi-loss collaborative optimization method is used for the ice segmentation image to generate a high-precision ice segmentation image. The final segmentation map is shown in Figure 2. Figure 5 As shown;

[0166] Finally, according to step 6, the segmentation feature map is optimized through the multimodal fusion network, and the icing type recognition is completed by combining the Softmax activation function. At the same time, the dynamic loss weight strategy is used to optimize the icing type recognition effect. Subsequently, according to step 7, the segmentation results and type recognition results are input into the multi-task hybrid attention network (MTHAnet), and the full connection layer and attention mechanism are combined to complete the preliminary classification of the ice thickness level. The classification accuracy is further optimized through the dynamic weighted loss function to obtain the final thickness level classification result.

[0167] In addition, advanced classification models such as VGG16, MobileNetV3_small, and ResNeXt101 were used to train and test on the ice type recognition dataset. The accuracy curves of these models on the validation set were recorded and compared with Swin-DeepSeg. The experimental results are shown in the figure below. Figure 7 As shown, it can be seen that the verification accuracy of Swin-DeepSeg can be stabilized at around 87%, which is higher than the other models, and the model stability of Swin-DeepSeg is also better than the other models.

[0168] Example 3

[0169] An electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the power line icing level classification method in embodiment 1, and the processor is configured to execute the program stored in the memory.

[0170] Example 4

[0171] A storage medium stores a computer program, which, when executed by a processor, executes the steps of the method for classifying ice coverage levels of power lines in embodiment 1.

[0172] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for classifying ice coverage levels of power lines, characterized in that: The following steps are involved: Performing preprocessing operations on the original ice-covered image to obtain a standard ice-covered image; The DConvLSTM spatiotemporal feature extraction model is used to extract meteorological spatiotemporal features. The global and local features of the standard ice-covered image are extracted using Swin Transformer and ResNet, respectively. Spatial pyramid pooling is then used to fuse the global and local features of the image to obtain multi-scale features. Terrain features are then introduced, and the meteorological spatiotemporal features, terrain features, and multi-scale features are weighted and summed to obtain a spatiotemporal multimodal feature map. The spatiotemporal multimodal feature map is input into the encoder composed of enhanced multi-head self-attention and deformable adaptive block, and the convolutional block attention module is combined to achieve feature weighted enhancement to obtain the enhanced encoded feature map; A lightweight deconvolution module is used to restore image resolution. Residual connections, lightweight convolution, fused convolution, and an efficient channel attention mechanism are combined to optimize feature expression. Finally, bilinear interpolation is used to refine image features. A multi-loss collaborative optimization method is used to generate high-precision ice segmentation maps for ice segmentation images. The spatiotemporal multimodal feature map and ice area segmentation map are input into the multimodal fusion network to extract and optimize the segmentation feature map. The ice type is identified through the Softmax activation function, and the recognition result is optimized through the dynamic loss weight. The ice segmentation results and ice type recognition results are input into the multi-task hybrid attention network, and the ice thickness level is preliminarily classified through the fully connected layer and attention mechanism. The dynamic weighted loss function is introduced to obtain the optimized ice thickness level classification, thereby obtaining the final ice thickness level result.

2. The method for classifying power line icing levels according to claim 1, characterized in that: The method for preprocessing the original ice-covered image to obtain a standard ice-covered image is as follows: obtaining the original ice-covered image, using an improved bilinear interpolation method and a Z-score detection method to process image missing values and outliers, performing wavelet denoising with soft and hard threshold fusion on the processed ice-covered image, using the CLAHE algorithm to perform enhanced adaptive histogram equalization processing, and obtaining a standard ice-covered image after normalization.

3. The method for classifying power line icing levels according to claim 1, characterized in that: The calculation formula of the spatiotemporal multimodal feature map is as follows: in, represents the spatiotemporal multimodal feature map, F ERA5 represents the spatiotemporal characteristics of meteorology, F DEM Indicates terrain features, F SPP represents multi-scale features, and α, β, and γ all represent weight parameters.

4. The method for classifying power line icing levels according to claim 1, characterized in that: The enhanced encoding feature map is represented as follows: Z enhanced =Z c (i,j,d)+Z s (i,j,d) Among them, Z enhanced Represents the enhanced encoding feature map, Z c (i, j, d) represents the result of channel attention optimization of the i-th row, j-th column, and d-th channel of the input image, Z s (i, j, d) represents the result of spatial attention optimization on the i-th row, j-th column, and d-th channel of the input image.

5. The method for classifying power line icing levels according to claim 1, characterized in that: The method of restoring image resolution using a lightweight deconvolution module is as follows: using 1×1 convolution to compress the number of channels of the enhanced encoded feature map, reducing the number of channels to half of the original number, using 2×2 transposed convolution to enlarge the spatial size of the feature map to twice the original size, and then using another 1×1 convolution to restore the number of channels to the number of channels of the input feature map, thereby obtaining a spatial high-resolution feature map.

6. The method for classifying power line icing levels according to claim 1, characterized in that: The total loss function of the generator in the multi-loss collaborative optimization method is expressed as follows: L G =λ1·L adv +λ2·L Dice +λ3·L Focal Among them, L G Denotes the total loss function of the generator, L adv represents the adversarial loss between the generator and the discriminator in the generative adversarial network framework, L Dice represents the Dice similarity loss between the generated image and the true label, L Focal Represents Focal loss, λ1, λ2, and λ3 represent the weight coefficients of the corresponding loss function.

7. The method for classifying power line icing levels according to claim 1, characterized in that: The method of inputting the spatiotemporal multimodal feature map and the ice area segmentation map into the multimodal fusion network to extract and optimize the segmentation feature map, realizing ice type recognition through the Softmax activation function, and optimizing the recognition result through the dynamic loss weight is as follows: The multimodal fusion network extracts and optimizes the initial ice segmentation map and performs classification output. The Softmax function is used to obtain the category probability distribution of each pixel. The cross entropy loss function is used to calculate the difference between the prediction and the true label, and dynamic loss weights are added to balance the influence of different categories.

8. The method for classifying power line icing levels according to claim 1, characterized in that: The method of inputting the ice segmentation results and ice type recognition results into the multi-task hybrid attention network, performing preliminary classification of ice thickness levels through the fully connected layer and attention mechanism, introducing a dynamic weighted loss function to obtain optimized ice thickness level classification, and thus obtaining the final ice thickness level result is as follows: The ice segmentation results and ice type recognition results are reduced in dimension through convolution operations, and the attention mechanism is applied in the channel dimension and spatial dimension respectively. The added together obtains the attention-enhanced fusion features, which are expressed as follows: F hybrid =F c +F s Among them, F hybrid represents the fusion feature after attention enhancement, F c represents the channel attention feature, F s Represents spatial attention characteristics; The fusion feature F after attention enhancement hybrid Using a fully connected layer and a dynamic weighted loss function, the output is an optimized ice thickness level prediction result, which is expressed as follows: Y=Softmax(FC(F hybrid )) Among them, FC represents the fully connected layer and Y represents the probability distribution of each category.

9. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the power line icing level classification method according to any one of claims 1 to 8, and the processor is configured to execute the program stored in the memory.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for classifying the ice coverage level of a power line according to any one of claims 1 to 8 are executed.

Citation Information

Cited By

  • Soil physicochemical property downscaling method and device, computer equipment and medium

    CN121072346A

  • Power transmission line inspection image denoising method and system based on SADNet-T network

    CN121788956A