Edge extraction method, device and equipment of ice-coated power line, medium and product
Through the method of combining multi-scale CNN and Transformer, the multi-scale local features of the ice-covered power line image are extracted, which solves the problem of low accuracy in extracting edge information of the ice-covered power line, realizes more accurate edge information generation, and enhances the robustness and adaptability of the model.
Patent Information
- Application Number
- CN202510360093.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, under extreme weather conditions, the extraction accuracy of the edge information of the ice-covered power line is low, resulting in inaccurate calculation of the ice-covered thickness, affecting the safety assessment and deicing strategy of the power line.
Multi-scale local features of ice-covered power line images are extracted through multi-scale CNN and combined with Transformer's encoder and decoder to generate more accurate edge information using a combination of global and local features.
It improves the accuracy of extracting edge information of ice-covered power lines, ensures overall continuity of edges and the accuracy of local details, enhances the model's adaptability to complex backgrounds, and avoids misjudgment of noise and local changes.
Smart Images

Figure CN120472184A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to image processing, and in particular to a method, device, equipment, medium and product for extracting the edge of ice-covered power lines. Background Art
[0002] Under extreme weather conditions, such as heavy snow, fog, or freezing rain, power lines can become covered in ice. Icing can significantly increase the weight and mechanical stress of power lines, so calculating the ice thickness on power lines is crucial.
[0003] Currently, calculating the ice thickness of power lines primarily involves inputting an image of ice-covered power lines into a Transformer, which then extracts edge information from the ice-covered power lines. Based on this extracted edge information, the ice thickness of the power lines is then calculated.
[0004] However, the accuracy of extracting edge information of existing ice-covered power lines is low. Summary of the Invention
[0005] The embodiments of the present application provide an edge extraction method, apparatus, device, medium, and product for ice-covered power lines, so as to improve the accuracy of edge information extraction of ice-covered power lines.
[0006] In a first aspect, an embodiment of the present application provides a method for edge extraction of ice-covered power lines, comprising:
[0007] Multi-scale CNN is used to extract features from the ice-covered power line image to obtain local features at multiple scales.
[0008] Fusing the multiple local features of different scales to generate a first feature;
[0009] Inputting the first feature into the encoder of the Transformer to obtain the second feature output by the encoder of the Transformer;
[0010] The second feature and the multiple local features of different scales are input into the decoder of the Transformer to obtain edge information of the ice-covered power line image to be processed output by the decoder of the Transformer.
[0011] In a possible implementation, the Transformer encoder includes a deformable convolution module, the deformable convolution module includes a first DFC, and inputting the first feature into the Transformer encoder to obtain a second feature output by the Transformer encoder includes:
[0012] Calculating a first offset according to the first feature using the first DFC;
[0013] Performing deformable sampling on each channel of the first feature according to the first offset using the first DFC to generate a third feature;
[0014] The third feature is processed by a multi-head attention mechanism and a feedforward network in the encoder of the Transformer to generate the second feature.
[0015] In a possible implementation manner, calculating the first offset according to the first feature using the first DFC includes:
[0016] Processing the first feature using multiple target convolution kernels with different dilation rates in the first DFC to generate a first initial offset corresponding to each target convolution kernel;
[0017] The first initial offsets corresponding to all target convolution kernels are fused to generate the first offset.
[0018] In a possible implementation, performing deformable sampling on each channel of the first feature according to the first offset using the first DFC to generate the third feature includes:
[0019] For each channel, superimpose the first offset on the initial sampling position of the standard convolution kernel through the first DFC to determine the sampling position;
[0020] If the sampling position is a non-negative integer, the feature value of the sampling position in the first feature is interpolated to generate the third feature.
[0021] In one possible implementation, the Transformer encoder and the Transformer decoder are connected via a skip connection module, where the skip connection module includes a channel attention submodule and a spatial attention submodule.
[0022] In a possible implementation, the skip connection module includes a second DFC, where the second DFC is used to calculate a second offset based on the second feature, and perform deformable sampling on each channel of the second feature based on the second offset to generate a processed second feature.
[0023] In a second aspect, an embodiment of the present application provides an edge extraction device for ice-covered power lines, comprising:
[0024] A feature extraction module is used to extract features from the ice-covered power line image to be processed using a multi-scale CNN to obtain local features at multiple scales;
[0025] A feature fusion module, configured to fuse the multiple local features of different scales to generate a first feature;
[0026] A first input module, configured to input the first feature into an encoder of a Transformer to obtain a second feature output by the encoder of the Transformer;
[0027] The second input module is used to input the second feature and the multiple local features of different scales into the decoder of the Transformer to obtain the edge information of the ice-covered power line image to be processed output by the decoder of the Transformer.
[0028] In one possible implementation, the Transformer encoder includes a deformable convolution module, the deformable convolution module includes a first DFC, and the Transformer encoder is specifically configured to:
[0029] Calculating a first offset according to the first feature using the first DFC;
[0030] Performing deformable sampling on each channel of the first feature according to the first offset using the first DFC to generate a third feature;
[0031] The third feature is processed by a multi-head attention mechanism and a feedforward network in the encoder of the Transformer to generate the second feature.
[0032] In one possible implementation, the Transformer encoder is specifically configured to:
[0033] Processing the first feature using multiple target convolution kernels with different dilation rates in the first DFC to generate a first initial offset corresponding to each target convolution kernel;
[0034] The first initial offsets corresponding to all target convolution kernels are fused to generate the first offset.
[0035] In one possible implementation, the Transformer encoder is specifically configured to:
[0036] For each channel, superimpose the first offset on the initial sampling position of the standard convolution kernel through the first DFC to determine the sampling position;
[0037] If the sampling position is a non-negative integer, the feature value of the sampling position in the first feature is interpolated to generate the third feature.
[0038] In one possible implementation, the Transformer encoder and the Transformer decoder are connected via a skip connection module, where the skip connection module includes a channel attention submodule and a spatial attention submodule.
[0039] In a possible implementation, the skip connection module includes a second DFC, where the second DFC is used to calculate a second offset based on the second feature, and perform deformable sampling on each channel of the second feature based on the second offset to generate a processed second feature.
[0040] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor;
[0041] The memory stores computer-executable instructions;
[0042] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.
[0043] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementation methods of the first aspect.
[0044] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.
[0045] The edge extraction method, device, equipment, medium and product of ice-covered power lines provided in the embodiments of the present application, the method jointly assists the Transformer decoder in generating edge information through the second feature (global feature) and local features of different scales, so that the Transformer decoder can simultaneously use the global context to determine the overall structure of the edge and ensure the overall continuity of the edge, while using local features of different scales to obtain fine features of the edge, and can repair edge breaks or discontinuous areas that may be ignored by the second feature to ensure the accuracy of the edge information. Furthermore, the local features can enhance the robustness of the Transformer to noise and local changes, avoid the excessive smoothing of local details by the second feature, and the second feature provides stable contextual information to avoid the misjudgment of the overall structure by the local feature. In the ice-covered scene, the ice-covered power line image can be better processed to generate more reliable edge information. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0047] Figure 1 Schematic diagram of the process of edge extraction method of ice-covered power lines provided in this application Figure 1 ;
[0048] Figure 2 Schematic diagram of the structure of the Transformer encoder provided in this application Figure 1 ;
[0049] Figure 3 Schematic diagram of the structure of the Transformer encoder provided in this application Figure 2 ;
[0050] Figure 4 Schematic diagram of the process of edge extraction method of ice-covered power lines provided in this application Figure 2 ;
[0051] Figure 5 Schematic diagram of the Transformer structure provided for this application Figure 1 ;
[0052] Figure 6 Schematic diagram of the Transformer structure provided for this application Figure 2 ;
[0053] Figure 7 A schematic diagram of the structure of the edge extraction device for ice-covered power lines provided in this application;
[0054] Figure 8 This is a schematic diagram of the structure of the electronic device provided in this application.
[0055] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0056] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0057] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0058] First, the terms used in this application are explained:
[0059] Transformer: A deep learning model architecture based on the self-attention mechanism. The core idea is to capture the global dependencies between different parts of the input data through the self-attention mechanism, thereby achieving parallel processing and efficient modeling of sequence data. The main components of the Transformer include the encoder (English: Encoder) and the decoder (English: Decoder). The encoder is composed of multiple identical layers stacked together, each of which contains two submodules: a multi-head self-attention mechanism and a feed-forward neural network. The self-attention mechanism dynamically assigns attention weights by calculating the correlation between each element in the input sequence and other elements, thereby capturing global contextual information. The decoder has a similar structure to the encoder, but additionally introduces a cross-attention mechanism to focus on the encoder output when generating the output.
[0060] Multi-Scale Convolutional Neural Network (CNN): A deep learning architecture that enhances the model's perception capabilities by extracting features at different scales. It uses multiple parallel convolution branches, each with a different-sized convolution kernel or receptive field to extract features at different scales.
[0061] Next, the application background of this application is explained:
[0062] Extracting the edges of power lines is a crucial task in the operation and maintenance of power systems. This edge information not only helps determine the precise location and shape of power lines but also provides essential data for subsequent fault detection, condition assessment, and safety warnings. For example, during power line inspections, images captured by drones or cameras require edge extraction technology to identify the outlines of power lines, thereby determining whether they are sagging, broken, or in contact with other objects.
[0063] Traditional methods for extracting power lines are mainly based on image processing technology, which uses the obvious contrast between power lines and background to determine the edge information of power lines.
[0064] Under extreme weather conditions, such as heavy snow, fog, or freezing rain, power lines can become covered in ice. Iced power lines significantly increase their weight and mechanical stress. If the ice thickness exceeds the designed load-bearing capacity of the power lines, it can cause power lines to break or power towers to collapse, leading to widespread power outages and even accidents. Therefore, calculating ice thickness is not only a key indicator for assessing the safety status of power lines, but also an important basis for developing de-icing strategies and preventative measures. By accurately calculating ice thickness, the ice coverage of power lines can be monitored in real time, allowing for timely de-icing measures to be implemented, avoiding power system failures caused by excessive ice, and ensuring the stability and security of the power supply.
[0065] However, under the above extreme weather conditions, the color of the ice-covered power lines is close to the background color. Using the above existing technology to extract the edges of the ice-covered power lines has the problem of low accuracy, which leads to low accuracy in the calculated ice thickness.
[0066] With the rapid development of machine learning, existing technologies have proposed that the edge information of ice-covered power lines can be extracted through Transformer. Specifically, the input ice-covered power line image needs to be preprocessed first, including grayscale, normalization, and resizing operations to eliminate noise and adapt to the model input. Subsequently, the ice-covered power line image is input into the encoder part of the Transformer. The encoder captures the global dependencies between different regions in the ice-covered power line image through the self-attention mechanism, such as the contrast relationship between the ice-covered power lines and the background, thereby extracting the overall contour features of the power lines. Then, these global features are passed to the decoder part. The decoder further refines the edge features in combination with local context information, and then generates a power line edge map.
[0067] However, while the Transformer effectively captures global features, its core is the self-attention mechanism, which extracts features by computing global dependencies between all pixels in the input image. While this mechanism effectively captures global information, it significantly increases computational complexity and memory consumption when processing high-resolution images, making it difficult for the model to accurately model local details. Furthermore, the self-attention mechanism tends to focus on prominent areas in the image (such as the overall outline of ice-covered power lines), while ignoring subtle edge variations or low-contrast local features (such as the blurred boundary between ice-covered power lines and the background). Furthermore, the Transformer's fixed position encoding may not accurately reflect the spatial relationships of local details when dealing with dynamic and complex backgrounds, further weakening its ability to capture local features. Therefore, in the task of extracting ice-covered power line edges, while the Transformer provides a global perspective, it performs relatively poorly in processing local details (such as the precise edges of power lines or subtle variations in ice cover), resulting in limited edge extraction precision and low extraction accuracy.
[0068] Based on the above technical problems, the technical concept of this application is as follows: Before processing the ice-covered power line image through the Transformer, the multi-scale local features of the ice-covered power line image are first extracted through a multi-scale CNN, and then the multi-scale local features are input into the Transformer, so that the Transformer can combine the multi-scale local features to capture global context information, enhance the model's adaptability to complex backgrounds, and then determine the edge information of the ice-covered power line image. By combining local features with global features, the edges of the ice-covered power lines can be refined, ensuring the clarity and accuracy of the final edge extraction.
[0069] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0070] Figure 1 Schematic diagram of the process of edge extraction method of ice-covered power lines provided in this application Figure 1 ,like Figure 1 As shown, the edge extraction method of the ice-covered power line includes:
[0071] S11. Feature extraction is performed on the ice-covered power line image to be processed through multi-scale CNN to obtain local features of multiple different scales.
[0072] The execution subject of the embodiments of the present application is an electronic device, which can be a terminal device, such as a laptop computer, a desktop computer, a tablet computer, etc., or a server. In actual applications, whether the electronic device is a terminal device or a server can be determined based on actual conditions and is not specifically limited to this.
[0073] In practical applications, an image acquisition device can be installed near power lines to capture power line images in real time. The device can transmit these images to electronic devices in real time or at a specified frequency. Multi-scale CNNs and Transformers are pre-deployed in the electronic devices to process the power line images and obtain edge information.
[0074] It should be understood that the power lines in the above-mentioned power line image can be in an ice-covered state or in an un-iced state. In practical applications, the ice-covered power line image can be processed according to the method described in the embodiment of the present application to obtain edge information of the ice-covered power line image. The power line image that is not in an ice-covered state can also be processed in the same manner according to the method described in the embodiment of the present application to obtain edge information of the power line image that is not in an ice-covered state.
[0075] The multi-scale CNN is composed of convolutional layers with different kernel sizes connected in parallel to extract local features of the ice-covered power line image at different scales. These convolutional layers work in parallel to capture local details and feature information at different scales.
[0076] It should be understood that a plurality refers to two or more.
[0077] Among them, the features involved in the embodiments of this application include information such as color, texture and edge.
[0078] Optionally, before performing feature extraction on the ice-covered power line image to be processed by the multi-scale CNN, the ice-covered power line image to be processed may also be preprocessed.
[0079] For example, the denoising module can be used to remove noise (such as Gaussian noise and salt-and-pepper noise) from the ice-covered power line image to improve image quality. The smoothing module can also be used to smooth the ice-covered power line image to smooth high-frequency noise and retain key structural information.
[0080] It should be understood that the local features of multiple different scales are multiple small blocks (English: patches) extracted from the ice-covered power line image to be processed, wherein each small block contains pixel information of a certain range.
[0081] S12: Fusing multiple local features of different scales to generate a first feature.
[0082] In practical applications, weighted summation or splicing can be used to fuse multiple local features of different scales, and the local features of multiple scales can be merged into a unified feature representation, namely the first feature.
[0083] S13. Input the first feature into the encoder of the Transformer to obtain the second feature output by the encoder of the Transformer.
[0084] In one possible implementation, the Transformer encoder can be Figure 2 The structure diagram shown is shown. Figure 2 Schematic diagram of the structure of the Transformer encoder provided in this application Figure 1 .like Figure 2 As shown in the figure, the Transformer encoder is composed of n identical layers stacked together, each of which consists of two main parts: a multi-head attention mechanism and a feedforward network. These two main parts are followed by a cross-layer and normalization operation to alleviate the gradient vanishing problem and accelerate training.
[0085] exist Figure 2 In the Transformer encoder structure shown, a multi-head attention mechanism is used to calculate the attention score for the first feature, which is then encoded into a new hidden state vector. This new hidden state vector is then linearly mapped through a feedforward network. Through n repeated iterations, the Transformer encoder continuously extracts higher-level feature representations, captures global context, enhances the model's adaptability to complex backgrounds, and generates the second feature.
[0086] In another possible implementation, Figure 3 Schematic diagram of the structure of the Transformer encoder provided in this application Figure 2 .like Figure 3 As shown, in Figure 2 Based on the Transformer, the encoder also includes a deformable convolution module.
[0087] The deformable convolution module includes a first deformable convolution layer (Deformable Convolutional Layer, DFC).
[0088] In this structure, the first DFC calculates the first offset based on the first feature. Then, the first DFC performs deformable sampling on each channel of the first feature based on the first offset to generate the third feature. Finally, the third feature is processed by the multi-head attention mechanism and feedforward network in the Transformer encoder to generate the second feature.
[0089] It should be understood that the encoder in Transformer is Figure 3 When the structure shown is Figure 4 The detailed explanation is given in the illustrated embodiment and will not be repeated here.
[0090] It should be understood that the second feature contains the global context information of the ice-covered power line image to be processed, provides a global understanding of the entire ice-covered power line image to be processed, and will provide support for subsequent local optimization.
[0091] In ice-covered scenes, ice-covered power lines have similar colors or textures to the background (such as trees and the sky). Local features at different scales can effectively represent detailed information in the ice-covered power line image to be processed (such as the edges of the power lines and the shape of the ice). By fusing multiple local features at different scales, the Transformer encoder processes the first fused feature to capture the relationship between all pixels in the ice-covered power line image to be processed, integrate global context information, and combine local features to suppress background interference, enhancing the Transformer's adaptability to complex backgrounds.
[0092] S14: Input the second feature and multiple local features of different scales into the decoder of the Transformer to obtain edge information of the ice-covered power line image to be processed output by the decoder of the Transformer.
[0093] In one possible implementation, the second feature and multiple local features at different scales are aligned to the same resolution through upsampling or downsampling. They are then concatenated or element-wise added to form a fused feature. The Transformer decoder then gradually extracts high-level features to generate precise edge information.
[0094] Edge information can be implemented in the form of an image. For example, the edge information can be a target image for edge annotation of an image of ice-covered power lines. The annotated edges are refined edges that are optimized using local features. This preserves the true edge contour, removes noise, and enhances detail. Edge information can also be implemented in the form of a coordinate file, which stores the coordinate values of each point scattered within the refined edge.
[0095] The embodiment of the present application provides an edge extraction method for ice-covered power lines. First, a multi-scale CNN is used to extract features from an ice-covered power line image to be processed to obtain multiple local features of different scales. Then, the multiple local features of different scales are fused to generate a first feature. Then, the first feature is input into the encoder of the Transformer to obtain a second feature output by the encoder of the Transformer. Finally, the second feature and the local features of multiple scales are input into the decoder of the Transformer to obtain edge information of the ice-covered power line image to be processed output by the decoder of the Transformer. In this technical solution, the second feature provides the overall structural information of the ice-covered power line image to be processed, while the local features of different scales provide the detailed information of different-sized areas in the ice-covered power line image to be processed. The second feature and the local features of different scales are used to assist the decoder of the Transformer in generating edge information. This allows the decoder of the Transformer to simultaneously utilize the global context to determine the overall structure of the edge and ensure the overall continuity of the edge. At the same time, the local features of different scales are used to obtain the fine features of the edge. This can repair the edge breaks or discontinuous areas that may be ignored by the second feature to ensure the accuracy of the edge information. Furthermore, local features can enhance the robustness of Transformer to noise and local changes, avoiding excessive smoothing of local details by the second feature. The second feature provides stable contextual information, avoiding misjudgment of the overall structure by local features, and can better process ice-covered power line images in ice-covered scenes, generating more reliable edge information.
[0096] exist Figure 3 Based on the encoder structure of the Transformer shown in the figure, the implementation process of S13 is explained in detail.
[0097] Figure 4 Schematic diagram of the process of edge extraction method of ice-covered power lines provided in this application Figure 2 ,like Figure 4 As shown, S13 can be implemented by the following steps:
[0098] S41 . Calculate a first offset according to a first feature using a first DFC.
[0099] In a possible implementation, the first feature may be processed by an offset generation network (English: Offset Generation) in the first DFC to calculate the first offset.
[0100] In this implementation, the offset generation network includes a lightweight convolutional layer (such as 1×1Conv), and the first offset is calculated using the following formula:
[0101] Δp=Conv 1×1 (X)
[0102] Wherein, Δp is the first offset, and X is the first feature.
[0103] In another possible implementation, the first feature is processed using multiple target convolution kernels with different dilation rates in the first DFC to generate a first initial offset corresponding to each target convolution kernel. The first initial offsets corresponding to all target convolution kernels are then fused to generate the first offset.
[0104] The target convolution kernel refers to the convolution kernel used to calculate the offset in the first DFC.
[0105] The first offset can be calculated using the following formula:
[0106]
[0107] Among them, d is used to represent the number of the target convolution kernel, and D is used to represent the number of target convolution kernels.
[0108] It should be understood that the dilation rate controls the spacing of sampling points in the target convolution kernel. A larger dilation rate can expand the receptive field of the convolution kernel and capture features across a wider range. In the first DFC, multiple target convolution kernels with different dilation rates (e.g., 1, 2, and 3) are used. Each convolution kernel corresponds to a different receptive field, capable of capturing features at different scales.
[0109] In the specific calculation, all target convolution kernels process the first feature in parallel, and then obtain the processing results of all target convolution kernels, that is, the first initial offset Afterwards, the first initial offsets corresponding to all target convolution kernels are fused by weighted averaging or other methods to obtain the first offset.
[0110] In this implementation, since the target convolution kernels with different expansion rates capture different feature scales, the fused offset can integrate multi-scale information, making the target convolution kernel better adapt to the shape and structure of the power line edge, thereby improving the performance and robustness of the model.
[0111] S42 . Perform deformable sampling on each channel of the first feature according to the first offset using the first DFC to generate a third feature.
[0112] In one possible implementation, for each channel, the initial sampling position of the standard convolution kernel is superimposed with the first offset through the first DFC to determine the sampling position. If the sampling position is a non-negative integer, the eigenvalues of the sampling position in the first feature are interpolated to generate the third feature.
[0113] The sampling position can be calculated using the following formula:
[0114] p ′ =Δp+p
[0115] Among them, p ′ is the sampling position, and p is the initial sampling position of the standard convolution kernel.
[0116] It should be understood that in standard convolution, the initial sampling position of the convolution kernel is fixed. For example, for a 3×3 convolution kernel, its initial sampling position can be expressed as:
[0117] p∈{(-1,-1),(-1,0),(-1,1),(0,-1),(0,0),(0,1),(1,-1),(1,0),(1,1)}
[0118] Optionally, the deformable convolution module also includes a depthwise separable convolution layer (DWC) and a maximum pooling layer (MP). The first DFC performs deformable sampling on each channel of the first feature according to the first offset to generate a third feature; then, the third feature is subjected to a depthwise separable convolution operation by the DWC, and then the output of the DWC is downsampled by the MP to extract the maximum value of the local area to obtain the processed third feature. The processed third feature is input into the multi-head attention mechanism and the feedforward network in the Transformer encoder to generate the second feature.
[0119] S43. The third feature is processed through the multi-head attention mechanism in the Transformer encoder and the feedforward network to generate the second feature.
[0120] The multi-head attention mechanism calculates the attention score of the third feature, encodes the calculated attention score into a new hidden state vector, and linearly maps this new hidden state vector through the feedforward network. Through n repeated iterations, the Transformer encoder continuously extracts higher-level feature representations, captures global context information, enhances the model's adaptability to complex backgrounds, and generates the second feature.
[0121] In the above embodiment, a deformable receptive field is used to overcome the limitations of conventional convolution. The DFC receptive field can adjust its shape, size, and even direction based on the target feature (power line edge). An additional offset field is used to adjust the specific position of the sampling points within the receptive field, thereby better capturing the details of the ice edge.
[0122] Figure 5 Schematic diagram of the Transformer structure provided for this application Figure 1 .like Figure 5 As shown in Figure 1, the encoder of Transformer and the decoder of Transformer are connected through a skip connection module, where the skip connection module includes a channel attention submodule and a spatial attention submodule.
[0123] It should be understood that the dual attention mechanism, which focuses on the channel dimension and spatial dimension of features respectively through the channel attention submodule and the spatial attention submodule, can more accurately capture important features in the processed ice-covered power line image, improve the accuracy of edge extraction, and at the same time strengthen the weight of the edge part in the processed ice-covered power line image, reduce noise interference, and improve the clarity of edge extraction.
[0124] Figure 6 Schematic diagram of the Transformer structure provided for this application Figure 2 .like Figure 6 As shown, the skip connection module includes a second DFC, which is used to calculate a second offset according to the second feature, and perform deformable sampling on each channel of the second feature according to the second offset to generate a processed second feature.
[0125] It should be understood that the processing of the second DFC can be similar to that of the first DFC, and the processing flow and principles of the two are the same. That is, after generating the processed second features, it is necessary to input the processed second features and multiple local features at different scales into the Transformer decoder to obtain the edge information of the ice-covered power line image to be processed, which is output by the Transformer decoder.
[0126] It should be understood that the second DFC can adaptively fuse low-level details (such as ice crystal texture) with high-level semantics (such as wire outline) to reduce edge blur during the upsampling process.
[0127] It should be understood that before executing the edge extraction method of ice-covered power lines through Transformer, model training is required in advance to obtain Transformer.
[0128] In a specific implementation, the execution subject of the ice-covered power line edge extraction method can also be an electronic device with processing capabilities, such as a terminal or server. It should be understood that the electronic device for the ice-covered power line edge extraction method and the electronic device for performing the model training process can be the same device or different devices.
[0129] During the training process, the model can be trained using a training set, a validation set, and a test set to obtain the Transformer to be optimized. The training set, validation set, and test set all include multiple sample ice-covered power line images and annotation information for each ice-covered power line image. The annotation information is used to describe the power line edges in the ice-covered power line images.
[0130] After obtaining the Transformer to be optimized, it can be supplemented with training using a Generative Adversarial Network (GAN) to optimize it, generating a Transformer for edge extraction of ice-covered power lines. Specifically, the generator evaluates the edge information generated by the Transformer to be optimized, guiding the generator to optimize this edge information, and thus optimizing the Transformer. This adversarial training method can effectively improve the detail and realism of edge extraction.
[0131] Furthermore, a directional convolutional pyramid (DCP) can be embedded in the GAN generator to separate the generation paths of horizontal, vertical, and diagonal edges, thereby enhancing the modeling capability of power line sag morphology.
[0132] Figure 7 This is a schematic diagram of the structure of the edge extraction device for ice-covered power lines provided in this application, as shown in FIG. Figure 7 As shown, the edge extraction device 70 of the ice-covered power line provided in this embodiment includes:
[0133] The feature extraction module 71 is used to extract features from the ice-covered power line image to be processed through a multi-scale CNN to obtain local features of multiple different scales.
[0134] The feature fusion module 72 is used to fuse multiple local features of different scales to generate a first feature.
[0135] The first input module 73 is used to input the first feature into the encoder of the Transformer to obtain the second feature output by the encoder of the Transformer.
[0136] The second input module 74 is configured to input the second feature and a plurality of local features of different scales into the decoder of the Transformer to obtain edge information of the ice-covered power line image to be processed output by the decoder of the Transformer.
[0137] In one possible implementation, the Transformer encoder includes a deformable convolution module, which includes a first DFC. The Transformer encoder is specifically configured to:
[0138] A first offset is calculated according to the first feature using the first DFC.
[0139] Each channel of the first feature is deformably sampled according to the first offset by the first DFC to generate a third feature.
[0140] The third feature is processed through the multi-head attention mechanism in the Transformer encoder and the feedforward network to generate the second feature.
[0141] In one possible implementation, the Transformer encoder is specifically configured to:
[0142] The first feature is processed by multiple target convolution kernels with different expansion rates in the first DFC to generate a first initial offset corresponding to each target convolution kernel.
[0143] The first initial offsets corresponding to all target convolution kernels are fused to generate the first offset.
[0144] In one possible implementation, the Transformer encoder is specifically configured to:
[0145] For each channel, the preset position of the second convolution kernel in the first DFC is superimposed with the first offset to determine the sampling position.
[0146] If the sampling position is a non-negative integer, the feature value of the sampling position in the first feature is interpolated to generate the third feature.
[0147] In one possible implementation, the Transformer encoder and the Transformer decoder are connected via a skip connection module, which includes a channel attention submodule and a spatial attention submodule.
[0148] In a possible implementation, the skip connection module includes a second DFC, which is used to calculate a second offset based on the second feature, and perform deformable sampling on each channel of the second feature based on the second offset to generate a processed second feature.
[0149] The edge extraction device for ice-covered power lines provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.
[0150] Figure 8 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 8 As shown, the electronic device 80 provided in this embodiment includes: at least one processor 801 and a memory 802. Optionally, the electronic device 80 further includes a communication component 803. The processor 801, the memory 802 and the communication component 803 are connected via a bus 804.
[0151] During the specific implementation process, at least one processor 801 executes the computer-executable instructions stored in the memory 802, so that the at least one processor 801 performs the above method.
[0152] The specific implementation process of the processor 801 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.
[0153] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly implemented by a hardware processor or implemented by a combination of hardware and software modules in the processor.
[0154] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.
[0155] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be classified into address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.
[0156] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0157] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.
[0158] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0159] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in a device as discrete components.
[0160] The division of units is merely a logical functional division; actual implementations may employ alternative divisions, such as combining or integrating multiple units or components into another system, or omitting or disabling certain features. Furthermore, any direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units, either through an interface, electrical, mechanical, or other means.
[0161] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0162] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0163] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0164] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0165] Finally, it should be noted that those skilled in the art will readily identify other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The present invention is not limited to the precise structure described above and illustrated in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof. The scope of the present invention is limited solely by the appended claims.
Claims
1. A method for extracting the edge of ice-covered power lines, characterized in that: include: The multi-scale convolutional neural network (CNN) is used to extract features from the ice-covered power line image to obtain local features at multiple scales. Fusing the multiple local features of different scales to generate a first feature; Inputting the first feature into the encoder of the Transformer to obtain the second feature output by the encoder of the Transformer; The second feature and the multiple local features of different scales are input into the decoder of the Transformer to obtain edge information of the ice-covered power line image to be processed output by the decoder of the Transformer.
2. The method according to claim 1, characterized in that The Transformer encoder includes a deformable convolution module, the deformable convolution module includes a first DFC, and inputting the first feature into the Transformer encoder to obtain a second feature output by the Transformer encoder includes: Calculating a first offset according to the first feature using the first DFC; Performing deformable sampling on each channel of the first feature according to the first offset using the first DFC to generate a third feature; The third feature is processed by a multi-head attention mechanism and a feedforward network in the encoder of the Transformer to generate the second feature.
3. The method according to claim 2, characterized in that Calculating a first offset according to the first feature by using the first DFC includes: Processing the first feature using multiple target convolution kernels with different dilation rates in the first DFC to generate a first initial offset corresponding to each target convolution kernel; The first initial offsets corresponding to all target convolution kernels are fused to generate the first offset.
4. The method according to claim 2 or 3, characterized in that The performing deformable sampling on each channel of the first feature according to the first offset by using the first DFC to generate a third feature includes: For each channel, superimpose the first offset on the initial sampling position of the standard convolution kernel using the first DFC to determine the sampling position; If the sampling position is a non-negative integer, the feature value of the sampling position in the first feature is interpolated to generate the third feature.
5. The method according to any one of claims 1 to 3, characterized in that The Transformer encoder and the Transformer decoder are connected via a skip connection module, which includes a channel attention submodule and a spatial attention submodule.
6. The method according to claim 5, characterized in that The skip connection module includes a second DFC, where the second DFC is used to calculate a second offset according to the second feature, and perform deformable sampling on each channel of the second feature according to the second offset to generate a processed second feature.
7. An edge extraction device for ice-covered power lines, characterized in that: include: A feature extraction module is used to extract features from the ice-covered power line image to be processed using a multi-scale convolutional neural network (CNN) to obtain local features at multiple scales. A feature fusion module, configured to fuse the multiple local features of different scales to generate a first feature; A first input module, configured to input the first feature into an encoder of a Transformer and obtain a second feature output by the encoder of the Transformer; The second input module is used to input the second feature and the multiple local features of different scales into the decoder of the Transformer to obtain the edge information of the ice-covered power line image to be processed output by the decoder of the Transformer.
8. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.
10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 6 when the computer program is executed by a processor.