An integrated restoration method for aerial degraded images of power transmission lines

CN122550419APending Publication Date: 2026-08-11NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-27
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]本发明所要解决的技术问题是:针对输电线路无人机航拍图像在雾、雨、雪等复杂气象条件下易出现对比度下降、纹理细节缺失、线状结构边缘断裂以及雨雪伪影残留,且现有方法在多退化耦合、空间非均匀退化以及多尺度特征传递过程中易引入噪声传播、复原结果不稳定、细节恢复不足的缺陷,提出一种输电线路航拍退化图像的一体复原方法,在端到端框架下同时实现多复杂气象下的巡检图像退化抑制与结构细节保真,提升复原图像在电力巡检中的可辨识性与下游缺陷识别的可靠性

Benefits of technology

[0058]1、本发明为雾、雨、雪三类退化分量独立配置卷积权重、偏置项及自适应激活斜率,可针对性适配单一/混合退化场景,突破传统方法难以兼顾各类气象巡检图像退化的技术瓶颈,实现复杂气象下图像退化的稳定处理。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550419A_ABST
    Figure CN122550419A_ABST
Patent Text Reader

Abstract

The application discloses a kind of transmission line aerial photograph degradation image integrated restoration method, comprising: first, the degraded image is normalized to obtain network input tensor, multi-resolution features are extracted by multi-scale coding network, feature expression is enhanced by mixed window attention operator, then global context aggregation is completed by horizontal / vertical long-range aggregation and local aggregation branch, subsequently, cross-scale interaction injection and confidence gating fusion are generated to generate restoration features, output restoration image and degradation component and construct degradation image estimate, finally, model end-to-end training is completed based on multi-loss weighted joint objective function.The application can accurately adapt to the geometric characteristics of transmission line, effectively suppress the restoration artifact, while considering the restoration accuracy and real-time, significantly improve the clarity and detail fidelity of inspection image under complex weather, can provide high-quality image basis for downstream transmission line defect detection, suitable for unmanned aerial vehicle power inspection scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent inspection technology for power systems, specifically to an integrated method for restoring degraded aerial images of transmission lines. Background Technology

[0002] As a crucial component of the power system, the safe and stable operation of transmission lines relies on continuous inspection of key components such as towers, conductors, insulator strings, and hardware. With the widespread adoption of drone-based aerial inspections, the efficiency and coverage of inspection operations have significantly improved, and these drones are gradually being integrated with visual algorithms such as target detection and defect identification to achieve intelligent operation and maintenance. However, in real-world engineering scenarios, inspection tasks often need to be carried out under non-extreme weather conditions such as dense fog, light rain, or light snow. The acquired aerial images are easily affected by various meteorological factors, resulting in problems such as decreased contrast, color shift, texture blurring, and obstruction by rain streaks / snow particles. This leads to discontinuities at the edges of slender structures like conductors and loss of detail in minute components, thus affecting the accuracy and robustness of manual interpretation and downstream intelligent recognition algorithms.

[0003] In existing image restoration methods, early solutions often modeled each degradation type separately, such as dedicated models for defogging, deraining, or desnowing. While these methods are effective in specific scenarios, they typically struggle to adapt to combined variations of different weather degradations within a unified framework. When different degradation modes are spatially non-uniformly distributed, the restoration process is also prone to producing local darkening, color distortion, or residual artifacts. On the other hand, feature extraction methods based on local receptive fields are limited in their ability to capture large-scale non-uniform degradation (such as clump-like fog). While improved methods using attention mechanisms enhance long-range dependency modeling capabilities, the fixed-shape window division often fails to take into account the geometric features of "slender linear targets" and "directional rain and snow noise" in transmission line scenarios, easily leading to problems such as edge aliasing and texture discontinuity.

[0004] Furthermore, while multi-scale encoding and decoding structures can expand the receptive field and improve computational efficiency through downsampling, the downsampling process can easily lead to the smoothing or diffusion of small target features such as insulators at deeper layers. Simultaneously, if the cross-layer feature transfer between the encoder and decoder lacks an effective filtering mechanism, unsuppressed background noise may be introduced into the decoding process, resulting in edge artifacts or ghosting in the restored image. Since power transmission inspection images simultaneously contain a large background area and fine textures of power components, how to balance global degradation suppression and local structure fidelity in a unified model, and how to suppress noise propagation and enhance detail recovery during multi-scale feature interactions, remain pressing technical problems to be solved in this field. Summary of the Invention

[0005] The technical problem this invention aims to solve is that aerial images of power transmission lines captured by drones under complex weather conditions such as fog, rain, and snow are prone to decreased contrast, loss of texture details, broken edges of linear structures, and residual rain and snow artifacts. Furthermore, existing methods suffer from drawbacks such as noise propagation, unstable restoration results, and insufficient detail recovery during multi-degradation coupling, spatially non-uniform degradation, and multi-scale feature transfer. This invention proposes an integrated restoration method for degraded aerial images of power transmission lines. Within an end-to-end framework, it simultaneously achieves degradation suppression and structural detail preservation in inspection images under various complex weather conditions, improving the recognizability of the restored images in power line inspections and the reliability of downstream defect identification. To address these problems, this invention employs the following technical solution:

[0006] First, this invention proposes an integrated restoration method for degraded aerial images of power transmission lines, comprising the following steps:

[0007] S1. Obtain the degraded aerial image of the transmission line to be restored and perform normalization processing to obtain the network input tensor;

[0008] S2. Construct an end-to-end multi-scale encoding and decoding network framework. Input the network input tensor into the multi-scale encoding network to obtain multi-scale encoding features corresponding to different spatial resolutions. The spatial resolution is reduced and channel mapping is completed between adjacent scale levels through downsampling operators.

[0009] S3. Perform feature enhancement transformation on the encoded features at each scale, and obtain enhanced features by using residual connection method;

[0010] S4. Perform multi-branch global context aggregation and gating fusion operations on the enhanced features, and achieve anisotropic long-range information representation through dilated convolution to obtain global aggregated features;

[0011] S5. Input the global aggregated features into the multi-scale decoding network, obtain the restored features through the cross-scale feature optimization mechanism, output the restored image and degradation components, and construct the degradation image estimate.

[0012] S6. Based on the reconstruction loss, degradation consistency loss, and cross-scale consistency loss between the restored image and the real clear image, construct a multi-loss weighted joint objective function adapted to the integrated restoration of aerial images of power transmission lines and perform end-to-end training; input the degraded image into the trained multi-scale encoding and decoding network and output the restored image.

[0013] Preferably, in step S2, the downsampling operator is used to sample the first... Coding features at each scale Pixel rearrangement downsampling is performed, and the result is obtained through linear mapping. Coding features at each scale The processing formula is as follows:

[0014] ;

[0015] in, , Indicates the number of scale levels; Indicates step size Perform input features Neighborhood rearrangement operator: rearranges input features according to... The pixel blocks are divided into neighborhoods, and the pixels within each neighborhood are stacked and rearranged according to the channel dimension, ultimately making the spatial resolution [value missing]. And the number of channels becomes the original number of channels. times; Indicates the first The scale is used for the learnable linear transformation operator of channel mapping, and the calculation formula is as follows: ,in and These are the learnable weights and bias parameters.

[0016] Preferably, in step S3, the feature enhancement transformation employs a residual connection approach, sequentially performing layer normalization, hybrid window attention enhancement, and feedforward mapping enhancement on the encoded features to obtain enhanced features, specifically including:

[0017] For the Input features at each scale Layer normalization and hybrid window attention enhancement are performed, and intermediate enhanced features are obtained using residual connections. The calculation formula is as follows:

[0018] ;

[0019] in, Indicates the attention operator for mixed windows; Representation layer normalization operator;

[0020] For the Intermediate enhancement features at each scale Layer normalization and feedforward mapping enhancement are performed, and residual connections are used to obtain the output enhancement features. The calculation formula is as follows:

[0021] ;

[0022] in, This represents the feedforward mapping operator, which consists of a 2-layer fully connected network and GELU nonlinear activation.

[0023] Preferably, the hybrid window attention operator To accommodate the slender geometric features of transmission lines, cross-window correlation modeling is achieved through multi-window branch combination, and the processing procedure satisfies the following formula:

[0024] ;

[0025] ;

[0026] ;

[0027] in, , , , To enable learnable linear mapping operators, a 1×1 convolutional layer is used. It is a set of branches, and contains at least square window branches. Horizontal long strip window branches With vertical long strip window branches ; Subscript Note branches representing different window shapes; Indicates in branch Query after window division tensor; Indicates branch The corresponding keys in the lower window tensor, Indicates its transpose; Indicates branch The corresponding value in the lower window tensor; To branch The corresponding relative position offset term; For branches The adaptive scaling factor, which is always positive; Feature dimensions for each attention head; For branches Adaptive branch gating weights; This represents the Softmax normalization function.

[0028] Preferably, in step S4, the enhanced feature In the Global aggregate features are obtained at each scale through multi-branch global context aggregation and gating fusion operations. The processing procedure satisfies the following formula:

[0029] ;

[0030] ;

[0031] ;

[0032] in, For the first The scale output mapping operator is implemented using a 1×1 convolutional layer; This represents element-wise multiplication; , , The first Horizontal, vertical, and local branch weights at different scales; , , The first Learnable mapping operators for horizontal long-range aggregation, vertical long-range aggregation, and local aggregation at different scales are implemented using strip, strip, and square dilated convolutions, respectively, with each dilated convolution configured with a corresponding dilation rate. This is a global average pooling operator; For the first A learnable mapping operator that maps channel description vectors to branch weight vectors at scale is implemented using a 2-layer fully connected network; This represents the Sigmoid function.

[0033] Preferably, in step S5, the cross-scale feature optimization mechanism first achieves multi-scale feature complementarity optimization through cross-scale interactive injection, and then optimizes the injected scale features. The final restored features are obtained by performing confidence-gated fusion. Specifically, it includes the following steps:

[0034] S501, The cross-scale interactive injection of source scale features After scale alignment, weights are injected into the target scale features using pixel-level weight injection. In the process, the scale features after injection are obtained. The processing procedure satisfies the following formula:

[0035] ;

[0036] ;

[0037] in, Representation and Scale The set of source scale indices for interaction; Indicates source scale Features on the surface are upsampled / downsampled and then channel-aligned before being mapped to the target scale. The mapping operator uses interpolation for upsampling, pooling for downsampling, and 1×1 convolution for channel alignment. This represents an operator that aggregates input features in the spatial dimension and broadcasts them back to the spatial dimension; This represents a learnable mapping operator used to generate injected weights, consisting of convolutional layers and nonlinear activation functions; Represents the Sigmoid function; Weights are injected at the pixel level, with values ​​ranging from [0,1].

[0038] S502, The confidence-gated fusion for adjacent scales and Restoration features and Alignment is performed based on the confidence graph. Adaptively adjust the fusion weights to obtain the fused restored features. The processing procedure satisfies the following series of calculation formulas:

[0039]

[0040]

[0041] in, The upsampling alignment operator is implemented using interpolation, with the upsampling factor matched with the downsampling step size; To make the first The candidate decoding feature map at the same scale is a restored feature map operator at the same scale, consisting of convolutional layers, batch normalization layers and nonlinear activation functions; For the first The learnable mapping operator that generates confidence maps at scale consists of convolutional layers, global pooling operators, and channel attention operators.

[0042] Preferably, in step S5, the degradation component Includes degradation components such as fog, rain, and snow; constructs degradation image estimates. By restoring the image Superimposed degenerate components The corresponding pixel-domain residual is implemented and calculated as follows:

[0043] ;

[0044] in, To restore the image; For the first A degenerate component The fog, rain, and snow degradation components adopt an independent parameter configuration principle, with dedicated convolution weights, bias terms, and adaptive activation slopes set for each. The calculation relationship of each degradation component is as follows:

[0045] ;

[0046] in, For adaptive correction of linear units, slope These are learnable parameters; To make the first Each degenerate component is mapped to... The learnable mapping operator for pixel domain residuals of the same size consists of depthwise separable convolution, batch normalization layer, nonlinear activation function and linear mapping; It is a batch normalization operator; the fog, rain, and snow degradation components are configured with independent learnable convolution weights, bias terms, and adaptive activation slopes.

[0047] Preferably, in step S6, the multi-loss weighted joint objective function for the integrated restoration of aerial images of transmission lines is... Losses from reconstruction Degradation consistency loss and cross-scale consistency loss The calculation formula is as follows, which is jointly constructed through pixel-level weight fusion:

[0048] ;

[0049] in, , , These are the learnable scalar weights updated via backpropagation during training.

[0050] The weight hyperparameter for slender targets in transmission lines is set to [1.0, 2.0], and is used to strengthen the recovery loss constraints for slender targets such as towers and conductors.

[0051] The weighted hyperparameter for multiple degradation scenarios has a value range of [1.0, 1.5] and is used to strengthen the degradation consistency loss constraint in multiple degradation scenarios.

[0052] is the multi-scale feature weight hyperparameter, with a value range of [0.8, 1.2], used to balance the cross-scale consistency loss constraint of features at different scales;

[0053] The reconstruction loss was designed to address the structural characteristics of aerial images of power transmission lines, and to meet the requirements. , For the pixel mask of slender targets in power transmission lines, higher weights are assigned to pixels in core areas such as towers and conductors, while background pixels are assigned basic weights. The total number of pixels in the image. Smoothing constant ;

[0054] The degradation consistency loss is designed for scenarios with multiple degradation layers, such as fog, rain, and snow. This is a cross-scale consistency loss designed for multi-scale encoder-decoder networks. , All based on The core construction logic is adapted and adjusted according to the corresponding constraint objectives, and all of them include smoothing constants. .

[0055] Furthermore, the present invention proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the control method as described in the present invention.

[0056] Meanwhile, the present invention proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed, it implements the steps of the method described in the present invention.

[0057] The present invention, by adopting the above technical solution, has the following beneficial effects:

[0058] 1. This invention independently configures convolution weights, bias terms, and adaptive activation slopes for three types of degradation components: fog, rain, and snow. It can be specifically adapted to single / mixed degradation scenarios, breaking through the technical bottleneck of traditional methods that cannot take into account various types of meteorological inspection image degradation, and achieving stable processing of image degradation under complex weather conditions.

[0059] 2. This invention innovatively designs a hybrid window attention operator, which accurately matches the geometric features of slender structures such as transmission line conductors and towers by combining square, horizontal strip, and vertical strip multi-branch windows. Combined with a multi-branch global context aggregation mechanism, it effectively suppresses edge aliasing and artifact residue, significantly improves the detail fidelity of core components, and provides a clearer target shape for downstream defect detection.

[0060] 3. This invention adopts an integrated optimization mechanism that combines cross-scale interactive injection and confidence gating. This mechanism not only achieves complementary enhancement of features at different scales, but also adaptively adjusts the feature fusion weights through the confidence map, avoiding the diffusion of small target features and noise propagation caused by downsampling, thus balancing the contradiction between global degradation suppression and local detail restoration.

[0061] 4. This invention proposes a multi-loss weighted joint objective function, which introduces the target weight of the slender transmission line, the superimposed weight of multiple degradations, and the hyperparameter of multi-scale feature weights. The focus of loss constraints can be flexibly adjusted according to the needs of the inspection scenario. Combined with degradation consistency loss and cross-scale consistency loss, it effectively improves the convergence stability of model training and the reliability of restoration results. Attached Figure Description

[0062] Figure 1This is a flowchart of an integrated restoration method for degraded aerial images of power transmission lines proposed in this invention.

[0063] Figure 2 This is a schematic diagram of the feature enhancement transformation module in this invention.

[0064] Figure 3 This is a schematic diagram of the structure of the multi-branch global context aggregation and gating fusion operation module in this invention.

[0065] Figure 4 This is a schematic diagram of the structure of global aggregation features fused across scales in this invention.

[0066] Figure 5 This is a schematic diagram of the structure for constructing a degraded image estimate by combining confidence-gated fusion with degraded components in this invention. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be further described below with reference to the accompanying drawings. It should be understood that the following embodiments are only for explaining this invention and are not intended to limit the scope of protection of this invention. Equivalent substitutions or modifications made by those skilled in the art without departing from the spirit of this invention should all be covered within the scope of protection of this invention.

[0068] Example 1: This example is a specific implementation of an integrated restoration method for degraded aerial images of power transmission lines. (Refer to...) Figure 1 The method comprises six core steps: degraded image preprocessing (S10), multi-scale feature encoding (S20), hybrid window attention feature enhancement (S30), multi-branch global context aggregation (S40), cross-scale fusion and joint construction of degraded image estimates with degraded components (S50), and joint training and inference with multiple losses (S60). This embodiment strictly adheres to the principle of controlling a single variable in configuring the experimental environment to ensure the reproducibility of the technical solution. All core processes and formulas are consistent with the claims. Furthermore, by supplementing detailed parameters, operational details, environmental configurations, and comparative experiments of core hyperparameters, the implementation process of the method is fully presented.

[0069] (I) Basic experimental setup:

[0070] 1. Dataset configuration:

[0071] This embodiment employs a dual data source to construct a dataset of degraded aerial images of power transmission lines, totaling 200,000 samples. This dataset covers complex terrains including mountains, hills, plains, and valleys, and includes five typical degradation scenarios: dense fog, light rain, light snow, mixed rain and fog, and snow and fog overlay, ensuring comprehensive validation of the model's generalization ability. The first part consists of 150,000 samples (5 years of accumulated data from drone inspections by a provincial power grid company), and the second part is a publicly available dataset of power transmission line images (including artificially synthesized fog, rain, and snow degradation labels, 50,000 samples). All samples are divided into a training set (140,000 samples), a validation set (20,000 samples), and a test set (40,000 samples) in a 7:1:2 ratio. Each set of samples contains degraded images. +Realistic and clear images +Degradation type label. During data preprocessing, all images were standardized to a resolution of 2048×1080. Data augmentation strategies such as random horizontal flipping (probability 0.5), brightness fine-tuning by ±15%, and adding Gaussian noise (variance 0.002) were employed to avoid model overfitting. Degraded images underwent normalization, mapping pixel values ​​to... The interval is used to obtain the network input tensor. .

[0072] 2. Hardware and software environment:

[0073] In terms of hardware configuration, the CPU uses an Intel Ultra 7 (20 cores, 40 threads) and the GPU uses an NVIDIA RTX 4090 (24GB VRAM) to ensure efficient model training and inference. The software environment is built on the Ubuntu 22.04 LTS operating system, using the PyTorch 2.0.1 deep learning framework and Python 3.9.16 as the programming language, coupled with CUDA 11.8 and CuDNN 8.7 for accelerated computing. This system can be deployed on the edge computing module of a power line inspection drone (such as the NVIDIA Jetson series), or on a ground control station or cloud server.

[0074] 3. Basic training parameters:

[0075] The model was trained using the AdamW optimizer, with an initial learning rate set to... Weight decay Maximum training iterations: 300 epochs. The first 30 epochs are the warm-up phase (learning rate starts from...). linear growth to Cosine annealing decay (decay factor 0.96) is used for 30-250 epochs, and a fixed learning rate is used for 250-300 epochs. Fine-tuning. Five core dimensions were selected as evaluation metrics: Peak Signal-to-Noise Ratio (PSNR, measuring the sharpness of image restoration), Structural Similarity (SSIM, measuring the structural consistency between the restored image and the real image), and the improvement rate of object detection accuracy. Based on YOLOv8, the system detects key components of power transmission lines and compares the detection accuracy, artifact rate (AR, which is the percentage of pixels of artifacts such as conductor edge halo and rain and snow residue) and inference time (ms / frame, which measures real-time performance) before and after restoration, comprehensively quantifying the restoration effect.

[0076] (II) Detailed Implementation of Core Steps:

[0077] Figure 1 This is a flowchart of an integrated restoration method for degraded aerial images of power transmission lines proposed in this invention.

[0078] like Figure 1 As shown, the specific method for preprocessing the degraded image in step S10 is as follows:

[0079] Acquire aerial images of the degraded power transmission lines to be restored. Pixel-level normalization is performed on it, using the following formula:

[0080]

[0081] in, , Degraded images The minimum and maximum pixel values ​​are used to map the input image pixel values ​​to the specified values ​​through this normalization operation. The interval is used to obtain the network input tensor. This effectively improves the convergence speed of model training.

[0082] Furthermore, such as Figure 1 As shown, the multi-scale encoding in step S20 is specifically as follows:

[0083] Input tensor to the network Inputting into a multi-scale coding network yields multi-scale coding features corresponding to different spatial resolutions. ( In this embodiment, the preferred embodiment is... (i.e., 4 scale levels), spatial resolution reduction and channel mapping are achieved between adjacent scale levels through a patch-merging downsampling operator. The downsampling satisfies the following relationship:

[0084]

[0085] in, ; Indicates step size (Preferred) =2) Perform input features Neighborhood rearrangement operator: rearranges input features according to... The pixel blocks are divided into neighborhoods, and the pixels within each neighborhood are stacked and rearranged according to the channel dimension, ultimately making the spatial resolution [value missing]. And the number of channels becomes the original number of channels. times; Indicates the first The scale is used for the learnable linear transformation operator of channel mapping, and the calculation formula is as follows: ,in and For learnable weights and bias parameters, Dimensions , Dimensions The number of channels at each scale is configured as follows: =64、 =128、 =256、 =512.

[0086] Figure 2 This is a schematic diagram of the feature enhancement transformation module in this invention;

[0087] Furthermore, such as Figure 1 As shown, in step S30, the feature enhancement transformation affects the first... Enhancement processing is performed on the encoded features at each scale, specifically including:

[0088] For the The input features at each scale are subjected to layer normalization and blended window attention enhancement, and residual connections are used to obtain intermediate enhanced features. The calculation formula is as follows:

[0089]

[0090] in, Indicates scale index; Indicates the first Input features at each scale; This represents the intermediate enhanced feature obtained by adding the residual of the input feature after layer normalization and mixed window attention enhancement, and then adding it to the input feature. Indicates the attention operator for mixed windows;

[0091] The intermediate enhanced features are subjected to layer normalization and feedforward mapping enhancement, and residual connections are used to obtain the output enhanced features. The model structure is as follows: Figure 2 As shown, the calculation formula is as follows:

[0092]

[0093] in, This represents the output enhanced feature obtained by adding the intermediate enhanced feature to the residual after layer normalization and feedforward mapping enhancement, and then adding the intermediate enhanced feature to the residual. This represents the feedforward mapping operator, which consists of two fully connected layers and GELU nonlinear activation. The number of channels in the intermediate layer is four times the number of input channels, and the number of output channels is the same as the number of input channels (for example, when the number of input channels is 64, the number of intermediate layer channels is 256). The layer normalization operator is calculated using the following formula:

[0094]

[0095] in , These represent the mean and variance of the input features in the channel dimension, respectively. These are learnable parameters with a size equal to the number of input channels (initial values ​​are 1 and 0). Smoothing constant (preferred) ).

[0096] The hybrid window attention operator For any input tensor The following series of calculation formulas must be satisfied:

[0097]

[0098]

[0099]

[0100] in, Indicates the purpose of generating the output projection Learnable linear mapping operators, hybrid window attention operators The output is implemented using a 1×1 convolutional layer; A set of branches that contains at least square window branches. (Window size 7×7), horizontal long strip window branches (Window size 3×15) and vertical strip window branches (Window size 15×3); Subscript Note branches representing different window shapes; Indicates in branch Query after window division ( Tensor; Indicates branch The corresponding key in the lower window ( Tensor; Indicates its transpose; Indicates branch The corresponding value in the lower window ( Tensor; To branch The corresponding relative position bias term (randomly initialized, mean 0, variance 0.01); For branches The adaptive scaling factor is always positive to avoid instability caused by sign flipping during training. Feature dimensions for each attention head (preferred) ); For branches The adaptive branch gating weights satisfy and ; This represents the Softmax normalization function; , , Indicates the use of generating queries ,key and value The learnable linear mapping operators are all implemented using 1×1 convolutional layers.

[0101] To verify the rationality of the optimal window size combination related to adaptive branch gating weights, the principle of controlling a single variable was adopted. Other parameters were fixed as optimal values, and only the target parameter was changed. Performance evaluation was performed on the test set, and the comparative experimental results are shown in Table 1.

[0102] Table 1

[0103]

[0104] The experimental results show that when the window size is too small (5×5 / 3×11 / 11×3), the receptive field is insufficient, making it difficult to capture long-distance wire features, resulting in low restoration clarity and detection accuracy. When the window size is too large (9×9 / 3×19 / 19×3), background noise is easily introduced, and the wire edge artifact rate increases. The optimal combination of 7×7 / 3×15 / 15×3 can fully cover the wire features while suppressing background interference, and all core indicators show the best performance, verifying the rationality of this combination.

[0105] Figure 3 This is a schematic diagram of the structure of the multi-branch global context aggregation and gating fusion operation module in this invention;

[0106] Furthermore, such as Figure 1 As shown, step S40 enhances the features. Perform multi-branch global context aggregation and gating fusion operations, wherein the multi-branch includes at least a horizontal long-range aggregation branch, a vertical long-range aggregation branch, and a local aggregation branch, to obtain global aggregation features. Its network structure is as follows Figure 3As shown, the following series of calculation formulas are satisfied:

[0107]

[0108]

[0109]

[0110] in, For the first Scale enhancement features For the first Scale-based global aggregation features; For the first The scale output mapping operator is implemented using a 1×1 convolutional layer; This represents element-wise multiplication; , , The first Horizontal, vertical, and local branch weights at the scale and satisfying ; , , The first Learnable mapping operators for horizontal long-range aggregation, vertical long-range aggregation, and local aggregation at different scales are implemented using dilated convolutions of 3×15 (dilation rate 2), 15×3 (dilation rate 2), and 3×3 (dilation rate 1), respectively. Indicates the first Scale weights; For the global average pooling operator, Compressed into a channel-dimensional description vector; For the first A learnable mapping operator that maps channel description vectors to branch weight vectors at scale is implemented using a 2-layer fully connected network (input dimension). intermediate dimension (output dimension 3) This represents the Sigmoid function.

[0111] Figure 4 This is a schematic diagram of the structure of global aggregated features fused across scales in this invention; Figure 5 This is a schematic diagram of the structure for constructing a degraded image estimate by combining confidence-gated fusion with degraded components in this invention.

[0112] Furthermore, such as Figure 1 As shown, step S50 first achieves multi-scale feature complementarity optimization through cross-scale interactive injection, and then performs confidence-gated fusion based on the injected scale features to obtain the final restored features. Figure 4 The network structure diagram is shown below, and the specific steps include:

[0113] The cross-scale interactive injection satisfies the following series of calculation formulas:

[0114]

[0115]

[0116] in, Representation and Scale The source scale index set for interaction (e.g.) hour, ); Indicates source scale Features on the surface are upsampled or downsampled and then channel-aligned before being mapped to the target scale. The mapping operator, when When the source scale resolution is lower than the target scale, upsampling uses bilinear interpolation, and channel alignment is achieved through 1×1 convolution; when When the source scale resolution is higher than the target scale, downsampling is performed using average pooling (stride 2), and channel alignment is achieved through 1×1 convolution. This represents an operator that aggregates input features in the spatial dimension and broadcasts them back to the spatial dimension; This represents the learnable mapping operator used to generate the injected weights, and its calculation formula is: ,in For channel splicing operations, (Target scale characteristics) and (Source-scale features after broadcasting) spliced ​​along the channel dimension; Represents the Sigmoid function; Inject weights at the pixel level (range of values) ); The scale features are those after injection.

[0117] The network structure relationship satisfied by the confidence-gated fusion is as follows: Figure 5 As shown, the following series of calculation formulas are satisfied:

[0118]

[0119]

[0120] in, The restored features after fusion; For confidence plots; For the upsampling alignment operator, make the scale Feature alignment to scale Specifically, bilinear interpolation is used, with an upsampling factor of 1. ; Indicates the first Scale-based restoration features; Indicates the first Scale-based restoration features; For the first Candidate decoding features at scale; To make the first The candidate decoding feature map at the same scale is the restoration feature map operator at the same scale, which is implemented by 3×3 convolution + BatchNorm layer + ReLU activation, and the number of output channels is 3 (RGB image); For the first Learnable mapping operators that generate confidence maps at scales, satisfying the formula The channel attention operator satisfies the relation , and For learnable weight parameters ( Dimension 3× , Dimension 3). This represents the Sigmoid function.

[0121] Furthermore, such as Figure 1 As shown, in step S50, the estimated value of the degraded image is constructed from the restored image and the degraded component. The following relationship must be satisfied:

[0122]

[0123] in, To restore the image; For the first The calculation relationships of the degradation components of fog, rain, and snow are as follows:

[0124]

[0125] in, For adaptive correction of linear units, its slope For learnable parameters and (Fog degradation) Rain degradation Snow degradation ); To make the first Each degenerate component is mapped to... The learnable mapping operator for pixel-domain residuals of the same size is calculated using the following formula:

[0126] ;

[0127] The It is a 3×3 depthwise separable convolution. For batch normalization operators, The activation function is non-linear; the three degradation components—fog, rain, and snow—are each assigned independent convolutional weights. Bias terms and adaptive slope This enables targeted adaptation to different degradation modes.

[0128] Furthermore, such as Figure 1 As shown, the multi-loss joint training and inference described in step S60 are as follows:

[0129] Based on restored images With real and clear images Reconstruction losses between Degraded image estimation value With degraded images Degradation consistency loss and cross-scale consistency loss Construct a multi-loss weighted joint objective function and perform end-to-end training. The objective function satisfies:

[0130]

[0131] in, , , These are the learnable scalar weights updated via backpropagation during training.

[0132] The weight hyperparameter for slender targets of transmission lines (preferred value 1.5, range 1.0-2.0) is adjusted and optimized based on the actual data and restoration requirements of aerial images of transmission lines to strengthen the restoration loss constraints of slender targets such as towers and conductors;

[0133] The weighting hyperparameter for multiple degradation scenarios (preferred value 1.2, range 1.0-1.5) is adjusted and optimized based on the actual degradation type and superposition degree of fog, rain and snow in aerial images of transmission lines, to strengthen the degradation consistency loss constraint under multiple degradation scenarios;

[0134] The multi-scale feature weight hyperparameter (preferred value 1.0, range 0.8-1.2) is adjusted and optimized according to the scale level and feature resolution of the multi-scale encoder-decoder network to balance the cross-scale consistency loss constraint of features at different scales.

[0135] The reconstruction loss designed to address the structural characteristics of aerial images of transmission lines satisfies the formula:

[0136]

[0137] For the pixel mask of the slender target of the transmission line, the pixels in the core area such as the tower and the conductor are assigned a weight of 2.0, and the pixels in the background area are assigned a weight of 1.0.

[0138] The degradation consistency loss designed for scenarios with multiple degradation layers (fog, rain, snow) satisfies the following formula:

[0139]

[0140] The cross-scale consistency loss designed for multi-scale encoder-decoder networks satisfies the formula:

[0141]

[0142] , All based on The core construction logic is adapted and adjusted according to the corresponding constraint objectives, and all of them include smoothing constants. (Preferred) ), This represents the total number of pixels in the image.

[0143] The reasoning stage will degrade the image. Input the trained model and directly output the restored image. .

[0144] (III) Experimental Results and Comparative Analysis:

[0145] To ensure fairness in the comparison, this embodiment selects five mainstream state-of-the-art (SOTA) models that have performed exceptionally well in the field of image restoration in recent years as benchmarks. All algorithms use official open-source code and are retrained and tested under the same dataset, hardware environment, and evaluation metrics. Among them, Restormer is an image restoration model based on efficient Transformer, which performs well in tasks such as dehazing and deraining, and is known for its strong long-range dependency modeling ability; SwinIR combines SwinTransformer with image restoration technology and is a strong generalization model with significant advantages in multi-scale feature fusion; Uformer adopts a full Transformer architecture and effectively improves the ability to restore image details through a hierarchical attention mechanism; MIMO-UNet is an efficient multi-input multi-output network specifically designed for multi-degradation scenarios and has a fast inference speed; GRLNet is a multi-degradation restoration model for severe weather, which focuses on optimizing the processing logic for fog, rain, and snow overlay scenarios.

[0146] As shown in Table 2, the PSNR proposed in this invention is 2.1 dB higher than the current best-performing GRLNet, and the SSIM is 0.017 higher, indicating that the restored image is significantly better than the mainstream SOTA model in terms of overall clarity and structural fidelity, and the texture details of key components such as wires and insulators are more completely preserved. Compared to GRLNet, this invention achieves a 5.3 percentage point improvement, with an artifact rate (AR) of only 1.6%, far lower than other models (3.2%-4.5%). The core reason is that the hybrid window attention operator (adapting to the characteristics of slender conductors) and multi-branch global aggregation (suppressing edge artifacts) of this invention are specifically optimized for power transmission line scenarios. In contrast, general state-of-the-art models (such as Restormer and SwinIR) do not consider the special geometric features of power line inspection images, which easily leads to conductor edge aliasing and artifact residue. In terms of real-time performance and engineering adaptability, the inference time of this invention is 25ms / frame, which is not only far lower than Restormer, Uformer and other models, but also better than the highly efficient MIMO-UNet (28ms / frame), fully meeting the real-time processing requirements of UAV inspection. (Frame), balancing performance and efficiency, can be directly deployed in embedded terminals of drones.

[0147] Table 2

[0148]

[0149] Example 2: This example proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in this invention.

[0150] Example 3: This example proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed, it implements the steps of the method described in this invention.

[0151] It should be noted that the processing flow of embodiments 2-3 corresponds to the specific steps of the method provided in embodiment 1 of the present invention, and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in this embodiment can be found in the method provided in embodiment 1 of the present invention.

[0152] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0153] The specific implementation schemes described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific implementation schemes of the present invention and are not intended to limit the scope of the present invention. Any equivalent changes and modifications made by those skilled in the art without departing from the concept and principles of the present invention should fall within the scope of protection of the present invention.

Claims

1. A method for integrated restoration of degraded aerial images of power transmission lines, characterized in that, Includes the following steps: S1. Obtain the degraded aerial image of the transmission line to be restored and perform normalization processing to obtain the network input tensor; S2. Construct an end-to-end multi-scale encoding and decoding network framework. Input the network input tensor into the multi-scale encoding network to obtain multi-scale encoding features corresponding to different spatial resolutions. The spatial resolution is reduced and channel mapping is completed between adjacent scale levels through downsampling operators. S3. Perform feature enhancement transformation on the encoded features at each scale, and obtain enhanced features by using residual connection method; S4. Perform multi-branch global context aggregation and gating fusion operations on the enhanced features, and achieve anisotropic long-range information representation through dilated convolution to obtain global aggregated features; S5. Input the global aggregated features into the multi-scale decoding network, obtain the restored features through the cross-scale feature optimization mechanism, output the restored image and degradation components, and construct the degradation image estimate. S6. Based on the reconstruction loss, degradation consistency loss, and cross-scale consistency loss between the restored image and the real clear image, construct a multi-loss weighted joint objective function adapted to the integrated restoration of aerial images of power transmission lines and perform end-to-end training; input the degraded image into the trained multi-scale encoding and decoding network and output the restored image.

2. The method of claim 1, wherein, In step S2, the downsampling operator is used to sample the first... Coding features at each scale Pixel rearrangement downsampling is performed, and the result is obtained through linear mapping. Coding features at each scale The processing formula is as follows: ; in, , Indicates the number of scale levels; Indicates step size Perform input features Neighborhood rearrangement operator: rearranges input features according to... The pixel blocks are divided into neighborhoods, and the pixels within each neighborhood are stacked and rearranged according to the channel dimension, ultimately making the spatial resolution [value missing]. And the number of channels becomes the original number of channels. times; Indicates the first The scale is used for the learnable linear transformation operator of channel mapping, and the calculation formula is as follows: ,in and These are the learnable weights and bias parameters.

3. The method of claim 2, wherein, In step S3, the feature enhancement transformation employs a residual connection approach, sequentially performing layer normalization, hybrid window attention enhancement, and feedforward mapping enhancement on the encoded features to obtain enhanced features. Specifically, this includes: The input feature of the first scale is Layer normalization and mixed window attention enhancement are performed, and residual connection is adopted to obtain intermediate enhanced features The calculation formula is as follows:​ ; in, Indicates the attention operator for mixed windows; Representation layer normalization operator; For the first Intermediate enhancement features at each scale Layer normalization and feedforward mapping enhancement are performed, and residual connections are used to obtain the output enhancement features. The calculation formula is as follows: ; wherein, denotes a feed-forward mapping operator consisting of a 2-layer fully connected network with GELU non-linear activation.

4. The method of claim 3, wherein, The hybrid window attention operator To accommodate the slender geometric features of transmission lines, cross-window correlation modeling is achieved through multi-window branch combination, and the processing procedure satisfies the following formula: ; ; ; in, , , , To enable learnable linear mapping operators, a 1×1 convolutional layer is used. It is a set of branches, and contains at least square window branches. Horizontal long window branches With vertical long strip window branches ; Subscript Note branches representing different window shapes; Indicates in branch Query after window division tensor; Indicates branch The corresponding keys in the lower window tensor, Indicates its transpose; Indicates branch The corresponding value in the lower window tensor; To branch The corresponding relative position offset term; For branches The adaptive scaling factor, which is always positive; Feature dimensions for each attention head; For branches Adaptive branch gating weights; This represents the Softmax normalization function.

5. The method of claim 1, wherein, In step S4, the enhanced features In the Global aggregate features are obtained at each scale through multi-branch global context aggregation and gating fusion operations. The processing procedure satisfies the following formula: ; ; ; in, For the first The scale output mapping operator is implemented using a 1×1 convolutional layer; This represents element-wise multiplication; , , The first Horizontal, vertical, and local branch weights at different scales; , , The first Learnable mapping operators for horizontal long-range aggregation, vertical long-range aggregation, and local aggregation at different scales are implemented using strip, strip, and square dilated convolutions, respectively, with each dilated convolution configured with a corresponding dilation rate. It is a global average pooling operator; For the first A learnable mapping operator that maps channel description vectors to branch weight vectors at scale is implemented using a 2-layer fully connected network; This represents the Sigmoid function.

6. The method of claim 1, wherein, In step S5, the cross-scale feature optimization mechanism first achieves multi-scale feature complementarity optimization through cross-scale interactive injection, and then optimizes the injected scale features. The final restored features are obtained by performing confidence-gated fusion. Specifically, it includes the following steps: S501、the cross-scale interaction injection is performed on the source scale feature After the scale alignment, the target scale feature is injected in a pixel-level injection weight manner In this way, the injected scale feature is obtained The processing process satisfies the following formula: ; ; in, Representation and Scale The set of source scale indices for interaction; Indicates source scale Features on the surface are upsampled / downsampled and then channel-aligned before being mapped to the target scale. The mapping operator uses interpolation for upsampling, pooling for downsampling, and 1×1 convolution for channel alignment. This represents an operator that aggregates input features in the spatial dimension and broadcasts them back to the spatial dimension; This represents a learnable mapping operator used to generate injected weights, consisting of convolutional layers and nonlinear activation functions; Represents the Sigmoid function; Weights are injected at the pixel level, with values ​​ranging from [0,1]. S502, The confidence-gated fusion for adjacent scales and Restoration features and Alignment is performed based on the confidence graph. Adaptively adjust the fusion weights to obtain the fused restored features. The processing procedure satisfies the following series of calculation formulas: in, The upsampling alignment operator is implemented using interpolation, with the upsampling factor matched with the downsampling step size; To make the first The candidate decoding feature map at the same scale is a restored feature map operator at the same scale, consisting of convolutional layers, batch normalization layers and nonlinear activation functions; For the first The learnable mapping operator that generates confidence maps at scale consists of convolutional layers, global pooling operators, and channel attention operators.

7. The method of claim 1, wherein, In step S5, the degradation component Includes degradation components such as fog, rain, and snow; constructs degradation image estimates. By restoring the image Superimposed degenerate components The corresponding pixel-domain residual is implemented and calculated as follows: ; in, To restore the image; For the first A degenerate component The fog, rain, and snow degradation components adopt an independent parameter configuration principle, with dedicated convolution weights, bias terms, and adaptive activation slopes set for each. The calculation relationship of each degradation component is as follows: ; in, For adaptive correction of linear units, slope These are learnable parameters; To make the first Each degenerate component is mapped to... The learnable mapping operator for pixel domain residuals of the same size consists of depthwise separable convolution, batch normalization layer, nonlinear activation function and linear mapping; It is a batch normalization operator; the fog, rain, and snow degradation components are configured with independent learnable convolution weights, bias terms, and adaptive activation slopes.

8. The method of claim 1, wherein, In step S6, the multi-loss weighted joint objective function for the integrated restoration of aerial images of the adapted transmission lines is... Losses from reconstruction Degradation consistency loss and cross-scale consistency loss The calculation formula is as follows, which is jointly constructed through pixel-level weight fusion: ; wherein, , , are learnable scalar weights updated by backpropagation during training. For the transmission line slender target weight hyperparameter, the value range is [1.0, 2.0], which is used to strengthen the restoration loss constraint of slender targets such as towers and wires. For the multiple degradation superposition weight hyperparameter, the value range is [1.0, 1.5], which is used to strengthen the degradation consistency loss constraint in the multiple degradation scene; For the multi-scale feature weight hyperparameter, the value range is [0.8, 1.2], which is used to balance the cross-scale consistency loss constraint of different scale features; The reconstruction loss was designed to address the structural characteristics of aerial images of power transmission lines, and to meet the requirements. , For the pixel mask of slender targets in power transmission lines, higher weights are assigned to pixels in core areas such as towers and conductors, while background pixels are assigned basic weights. The total number of pixels in the image. Smoothing constant ; The degradation consistency loss is designed for scenarios with multiple degradation layers, such as fog, rain, and snow. This is a cross-scale consistency loss designed for multi-scale encoder-decoder networks. , All based on The core construction logic is adapted and adjusted according to the corresponding constraint objectives, and all of them include smoothing constants. .

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 8.

10. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed, it implements the steps of the method as described in any one of claims 1 to 8.