Power transmission line intelligent inspection method and system based on unmanned aerial vehicle and cloud side cooperation

By using a smart inspection method that combines drones and cloud-edge collaboration, and integrating static differential images and multi-view fusion detection, the problems of fixed viewpoints and detection delays in power transmission line inspections have been solved, enabling efficient and accurate defect identification and real-time response.

CN120976809AActive Publication Date: 2025-11-18WENZHOU ELECTRIC POWER CONSTR CO LTD

Patent Information

Application Number
CN202511494858.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Existing transmission line inspection methods suffer from problems such as fixed viewing angles, insufficient coverage, detection delays, and insufficient defect identification accuracy. In particular, they are difficult to achieve high-frequency, continuous inspections and real-time detection in complex environments.

Method used

By employing an intelligent inspection method that combines drones and cloud-edge collaboration, and integrating static sensing equipment with drone image acquisition, this method achieves efficient local defect detection and cloud-based multi-view analysis through static differential images and multi-view fusion detection, thereby improving the spatial coverage, temporal continuity, and accuracy of inspections.

Benefits of technology

It improves the robustness and real-time response capability of transmission line inspection, enhances the accuracy and efficiency of defect detection in complex environments, and ensures the efficiency and continuity of inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976809A_ABST
    Figure CN120976809A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power transmission line inspection, and provides an intelligent power transmission line inspection method and system based on unmanned aerial vehicle and cloud edge collaboration, and the method comprises the steps: generating a static differential image based on a static visual image of a target power transmission line, and carrying out the local defect detection to obtain a static detection result; when the static detection result meets the preset detection precision, the static detection result serves as a target inspection result, otherwise, the static visual image is uploaded to the cloud, and meanwhile, the unmanned aerial vehicle is triggered to execute an image supplementary collection task and uploads a collected dynamic inspection image data set to the cloud; and the cloud carries out defect identification analysis according to the first defect detection model used for carrying out multi-view feature fusion detection on the input image to obtain a target inspection result. According to the multi-view intelligent sensing architecture based on dynamic and static combination and side cloud multi-source cooperation, the space coverage capability and the time continuity of line inspection are effectively improved, and meanwhile, the efficiency and the accuracy of intelligent inspection of the power transmission line can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power transmission line inspection, in particular to a power transmission line intelligent inspection method and system based on cooperation of unmanned aerial vehicles and cloud edges. BACKGROUND

[0002] As a key channel for large-scale and long-distance power transmission in the power system, the operation state of the power transmission line is directly related to the stability of the power grid structure and the continuity of power supply. With the large-scale access of new energy, rapid growth of power demand and continuous expansion of power grid coverage, the distribution range of the power transmission line is becoming increasingly extensive, and the operation environment is becoming increasingly complex, resulting in a significant increase in defect occurrence frequency, which easily leads to power supply interruption, equipment damage and even regional power outage accidents, seriously threatening the safe operation of the power system. Therefore, efficient, accurate and sustainable state monitoring and defect identification of the power transmission line have become one of the key tasks of current power grid operation and management.

[0003] At present, the intelligent inspection of the power transmission line mainly adopts the way of fixedly installing static sensing devices or based on unmanned aerial vehicles and other dynamic devices to solve the application defects of manual inspection. However, although the existing intelligent inspection method has made significant progress in defect detection of the power transmission line, it still has application limitations: 1) the fixed perspective of the static sensing device makes it difficult to cover the entire line, and although the unmanned aerial vehicle has flexible collection capability, it is difficult to realize high-frequency and continuous inspection due to the endurance time and scheduling complexity, resulting in insufficient overall coverage and continuity; 2) uploading all the images collected by the static device or the unmanned aerial vehicle to the cloud for processing often introduces delay, affecting the real-time of detection; 3) the defect detection model relies on single perspective or isolated images for identification, and fails to fully exploit the complementarity and semantic association between multi-perspective information, and does not consider the difference in texture distribution characteristics of different target images, resulting in insufficient recognition accuracy of the complex structure and variable posture defect target of the power transmission line. SUMMARY

[0004] The purpose of the present application is to provide a power transmission line intelligent inspection method based on cooperation of unmanned aerial vehicles and cloud edges, which realizes local efficient defect detection based on the deployment of static sensing devices, adopts unmanned aerial vehicle emergency image supplement collection combined with cloud multi-perspective fusion detection analysis, and adopts a multi-perspective intelligent sensing architecture of dynamic and static combination and edge-cloud multi-source cooperation, which not only improves the spatial coverage capability and time continuity of line inspection, but also effectively improves the efficiency and accuracy of line inspection, thereby improving the robustness of intelligent inspection.

[0005] In order to achieve the above purpose, it is necessary to provide a power transmission line intelligent inspection method and system based on cooperation of unmanned aerial vehicles and cloud edges.

[0006] In a first aspect, an embodiment of the present application provides a power transmission line intelligent inspection method based on cooperation of an unmanned aerial vehicle and a cloud edge, the method comprising: acquiring a static visual image of a target power transmission line, generating a corresponding static difference image based on the static visual image, and performing local defect detection according to the static difference image to obtain a corresponding static detection result; when the static detection result meets a preset detection accuracy requirement, taking the static detection result as a target inspection result; in a case where the static detection result does not meet the preset detection accuracy requirement, uploading the static visual image to a cloud end in a compressed manner, triggering an unmanned aerial vehicle to perform an image supplementing task, and uploading a dynamic inspection image data set collected according to the image supplementing task to the cloud end in a compressed manner, so that the cloud end performs defect recognition analysis based on a first defect detection model pre-constructed according to the static visual image and the dynamic inspection image data set to obtain the target inspection result; the first defect detection model is used for multi-view feature fusion detection on an input image.

[0007] Further, the step of generating a corresponding static difference image based on the static visual image comprises: preprocessing the static visual image to obtain a preprocessed static image; acquiring a corresponding static view reference image based on a preset normal static visual image library according to an illumination condition of the preprocessed static image; generating a pixel difference image and a structure difference image according to the preprocessed static image and the static view reference image; performing weighted fusion on the pixel difference image and the structure difference image to obtain the static difference image.

[0008] Further, the step of performing local defect detection according to the static difference image to obtain a corresponding static detection result comprises: inputting the static difference image into a second defect detection model pre-constructed to perform defect recognition analysis, so as to obtain the static detection result; the second defect detection model is used for adaptive texture edge enhancement on the static difference image based on a texture complexity of the static difference image, and then performing semantic feature extraction and defect target detection in sequence.

[0009] Further, the second defect detection model comprises a dynamic routing module, a double-channel edge increasing module, a semantic feature extraction module and a detection head module connected in sequence; the double-channel edge increasing module comprises a first texture edge enhancement branch and a second texture edge enhancement branch connected in parallel. The dynamic routing module is configured to calculate a gray information entropy based on a proportion of pixel numbers of different gray levels in the static differential image, and input the static differential image into the first texture edge enhancement branch or the second texture edge enhancement branch in the double-channel edge increasing module according to a size relationship between the gray information entropy and a preset information entropy threshold. The double-channel edge increasing module is configured to perform adaptive texture edge enhancement on the static differential image based on a texture complexity of the static differential image, to obtain corresponding texture enhanced features. The semantic feature extraction module is configured to perform multi-scale semantic feature extraction on the texture enhanced features, to obtain corresponding multi-scale semantic features. The detection head module is configured to obtain the static detection result based on sequentially performing global average pooling and full connection processing on the multi-scale semantic features.

[0010] Further, the first texture edge enhancement branch comprises a sequentially connected cavity convolution layer, a sub-pixel convolution layer and a channel attention layer. The second texture edge enhancement branch comprises a sequentially connected depth convolution layer, an attention layer and a global pooling layer; the attention layer comprises a sequentially connected improved window attention layer and an improved shift window self-attention layer.

[0011] Further, the improved window attention layer and the improved shift window self-attention layer both modulate attention weights based on gradients of feature maps.

[0012] Further, the first defect detection model comprises a sequentially connected multi-view image conversion module, a multi-view feature extraction module, a feature enhancement fusion module and a defect target detection module. The multi-view image conversion module is configured to perform multi-view image conversion on the input image, to obtain a corresponding multi-view image set; the multi-view image set comprises a vertical top view image, a diagonal oblique view image and a side view image. The multi-view feature extraction module is configured to perform feature extraction on the vertical top view image based on a cavity spatial pyramid network, perform feature extraction on the diagonal oblique view image based on a deformable convolution network, and perform feature extraction on the side view image based on vertical stripe pooling combined with horizontal direction convolution, to obtain corresponding multi-view feature maps; the multi-view feature maps comprise a vertical top view feature map, a diagonal oblique view feature map and a side view feature map. The feature enhancement fusion module is configured to perform feature enhancement fusion on the multi-view feature maps and the input image based on a preset graph attention network and a multi-channel attention mechanism, to obtain corresponding multi-view fusion features. The defect target detection module is configured to perform target defect identification based on a preset neural network according to the multi-view fusion feature, and obtain a corresponding defect detection result.

[0013] Further, the feature enhancement fusion module comprises a graph convolution processing unit, a channel splicing unit, an attention fusion unit and a boundary enhancement unit connected in sequence. The graph convolution processing unit is configured to construct a full connection graph with the multi-view feature map and the input image as nodes, and perform update processing on the multi-view feature map and the input image according to an adjacency matrix of the full connection graph and the preset graph attention network, to obtain an updated multi-view feature map and an updated input image. The channel splicing unit is configured to perform channel splicing on the updated multi-view feature map and the updated input image, to obtain a corresponding joint feature tensor. The attention fusion unit is configured to calculate attention weight coefficients of each feature in the joint feature tensor based on a multi-channel attention mechanism of tensor modal product, and fuse the updated multi-view feature map and the updated input image based on the attention weight coefficients of each feature, to obtain an enhanced feature tensor. The boundary enhancement unit is configured to perform average pooling and maximum pooling on the enhanced feature tensor respectively, and perform convolution processing on the first pooled feature and the second pooled feature obtained correspondingly, and then fuse the enhanced feature tensor, to obtain the multi-view fusion feature.

[0014] Further, the cloud performs defect detection based on a first defect detection model pre-constructed according to the static visual image and the dynamic inspection image dataset, to obtain the target inspection result, and the steps comprise: Performing defect detection analysis on each frame image in the static visual image and the dynamic inspection image dataset based on the first defect detection model, to obtain a plurality of defect detection results; When all the defect detection results are defect-free, setting the target inspection result as normal inspection; When there is a defect target in all the defect detection results, obtaining a defect detection result with the highest confidence as the target inspection result.

[0015] In a second aspect, an embodiment of the present application provides a power transmission line intelligent inspection system based on cooperation of a UAV and a cloud edge, and the system comprises: A local detection module is configured to acquire a static visual image of a target power transmission line, generate a corresponding static differential image based on the static visual image, and perform local defect detection according to the static differential image, to obtain a corresponding static detection result. The precision analysis module is configured to, when the static detection result meets the preset detection precision requirement, take the static detection result as a target inspection result. The cloud detection module is configured to, when the static detection result does not meet the preset detection precision requirement, compress and upload the static visual image to the cloud, trigger the UAV to perform an image supplement task, and compress and upload a dynamic inspection image data set collected according to the image supplement task to the cloud, so that the cloud performs defect recognition analysis based on a first defect detection model pre-constructed for multi-view feature fusion detection of an input image according to the static visual image and the dynamic inspection image data set, to obtain the target inspection result.

[0016] The application provides a power transmission line intelligent inspection method and system based on UAV and cloud edge cooperation, which realizes obtaining a static visual image of a target power transmission line, generating a corresponding static difference image based on the static visual image, performing local defect detection according to the static difference image to obtain a corresponding static detection result, taking the static detection result as a target inspection result when the static detection result meets a preset detection precision requirement, and compressing and uploading the static visual image to the cloud when the static detection result does not meet the preset detection precision requirement, triggering the UAV to perform an image supplement task, and compressing and uploading a dynamic inspection image data set collected according to the image supplement task to the cloud, so that the cloud performs defect recognition analysis based on a first defect detection model pre-constructed for multi-view feature fusion detection of an input image according to the static visual image and the dynamic inspection image data set to obtain the target inspection result. Compared with the prior art, the power transmission line intelligent inspection method based on UAV and cloud edge cooperation realizes local efficient defect detection based on the deployment of static sensing devices, adopts a dynamic and static combination and edge-cloud multi-source cooperation multi-view intelligent sensing architecture of UAV emergency image supplement combined with cloud multi-view fusion detection analysis, can not only improve the spatial coverage capability and time continuity of line inspection, but also effectively improve the efficiency and accuracy of line inspection, and further improve the robustness and real-time response capability of intelligent inspection. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 is a flowchart of the power transmission line intelligent inspection based on UAV and cloud edge cooperation in the embodiment of the application; Figure 2 is a structural diagram of a second defect detection model in the embodiment of the application; Figure 3 is Figure 2 a structural diagram of a first texture edge enhancement branch in the embodiment of the application; Figure 4 is Figure 2A structural schematic diagram of a second texture edge enhancement branch; Figure 5 A structural schematic diagram of a first defect detection model in an embodiment of the present application; Figure 6 A structural schematic diagram of a power transmission line intelligent inspection system based on cooperation of unmanned aerial vehicles and cloud edges in an embodiment of the present application; Among them, the reference signs are: 1, a local detection module; 2, a precision analysis module; 3, a cloud detection module. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical scheme and beneficial effects of the present application clearer and more apparent, the present application will be further described in detail below in combination with the drawings and embodiments. Obviously, the following described embodiments are part of the embodiments of the present application, and are only used to illustrate the present application, but not to limit the scope of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.

[0019] In some embodiments, as shown in Figure 1 A power transmission line intelligent inspection method based on cooperation of unmanned aerial vehicles and cloud edges is provided, comprising the following steps: S11, obtaining a static visual image of a target power transmission line, generating a corresponding static difference image based on the static visual image, and performing local defect detection according to the static difference image to obtain a corresponding static detection result. The static visual image can be understood as a fixed perspective image at the deployment position obtained by the static perception device deployed at the key node of the power transmission line according to the preset sampling frequency. The static perception device can use an existing image acquisition device (such as a fixedly installed camera).

[0020] Considering that the actual static visual image acquisition process is easily disturbed by external light, which will affect the detection accuracy of defect detection based on the static visual image, in order to effectively resist external light interference and highlight structural defects (such as cracks, broken strands, etc.), and improve the detection performance of the local defect detection device, the embodiment preferably uses the AIBOX edge computing device to obtain the difference features of the static visual image based on the historical normal static perspective image as the analysis data for subsequent defect detection, in order to improve the robustness of light changes. It should be noted that the AIBOX edge computing device can realize real-time online monitoring of the images collected by the fixed camera, without the need to upload all the images collected by the fixed camera to the cloud for detection, so as to reduce the detection delay. Specifically, the step of generating a corresponding static difference image based on the static visual image comprises: The static visual image is preprocessed to obtain a preprocessed static image. The preprocessing may include size normalization, pixel normalization and noise removal of the static visual image to obtain a preprocessed static image with stable data quality. The specific processing procedure is implemented with reference to relevant existing technologies and will not be described in detail here.

[0021] Based on the illumination conditions of the preprocessed static image, a corresponding static viewpoint reference image is obtained from a preset normal static visual image library. Illumination conditions include brightness and contrast, which can be obtained through feature analysis of the preprocessed static image. For example, brightness can be extracted by calculating the average grayscale value of all pixels in the image, and contrast can be obtained by statistically analyzing the variance of the image's grayscale distribution. The preset normal static visual image library can be understood as a database containing images of normal transmission lines under different illumination conditions at different key points along the transmission line. It can be constructed and updated based on normal transmission line images collected by static sensing devices deployed at various key nodes along the transmission line. The static viewpoint reference image can be understood as an image obtained by matching and analyzing the preprocessed static image with normal transmission line images under the same illumination conditions in the preset normal static visual image library. This ensures that the obtained image has similar illumination conditions to the preprocessed static image, so that the subsequent static difference image reflects only the actual defects, not the illumination differences.

[0022] The actual matching analysis process for obtaining the static viewpoint reference image can be as follows: based on the preprocessed static image corresponding to the current time t... Lighting conditions, from a preset normal static visual image library Select the static viewpoint reference image whose histogram is most similar (highest histogram similarity). Preprocessing static images and each reference image in the preset normal static visual image library histogram similarity The correlation coefficient of the histogram is obtained using the following formula: in, For preprocessing still images In the Histogram values ​​at each gray level; To pre-define the k-th reference image in a normal static visual image library In the Histogram values ​​at each gray level, where N represents the total number of reference images in the preset normal static visual image library; and Preprocessed static images and reference image The histogram average.

[0023] According to the pre-processed static image and the static view reference image, a pixel difference image and a structure difference image are generated; wherein the pixel difference image can be understood as an image composed of the absolute value of the pixel difference at the corresponding position of the pre-processed static image and the static view reference image, which can directly reflect the gray level change between images. The structure difference image can be understood as a difference image constructed based on the structural similarity (SSIM) index in order to obtain a more robust difference image, considering that the pixel difference image is still sensitive to illumination changes and is easily disturbed by noise. In actual application, the local SSIM value of the pre-processed static image and the static view reference image can be calculated based on a preset sliding window, that is, a structural similarity matrix that can reflect the local structural consistency of the same position in the pre-processed static image and the static view reference image can be obtained, and then the obtained structural similarity matrix is up-sampled using a bicubic interpolation function, that is, a structure difference image that is aligned with the pixel difference image in spatial scale can be obtained. The calculation of the local SSIM value and the bicubic interpolation can refer to related prior art, which will not be described in detail here.

[0024] The pixel difference image and the structure difference image are weighted and fused to obtain the static difference image; wherein the static difference image can be represented as: wherein, and are the pre-processed static image and the corresponding static view reference image at time t, respectively; is the structural similarity matrix of and is a bicubic interpolation up-sampling function; is a dynamic weight coefficient; is the corresponding static difference image.

[0025] The static difference image obtained by fusing the structural similarity (SSIM) index on the basis of the traditional pixel difference in this embodiment can evaluate the image difference while effectively suppressing the influence of illumination, further enhancing the structural perception ability, and providing a reliable analysis basis for subsequent defect detection and recognition.

[0026] ​​The static differential image obtained through the above method steps can be used as basic data for defect target detection. To ensure the reliability of defect detection, the embodiment preferably performs fine-grained edge enhancement based on sufficient consideration of the texture distribution feature differences of different target images, so as to improve the precision and robustness of multi-type defect detection. Specifically, the step of performing local defect detection according to the static differential image to obtain a corresponding static detection result comprises: inputting the static differential image into a second defect detection model constructed in advance for defect recognition analysis to obtain the static detection result; the second defect detection model can be understood as a network model capable of analyzing the static differential image to obtain a static detection result including target bounding box position coordinates and detection confidence scores. Considering that the high-texture region in the image includes rich fine-grained texture information, a large receptive field and fine texture extraction capability are needed, while the low-texture region includes more structural dependency relationships, and a feature that can effectively suppress background noise and improve the expression efficiency of structural information is needed, the embodiment preferably adopts a network structure that can perform adaptive texture edge enhancement on the static differential image based on the texture complexity of the static differential image, and then sequentially perform semantic feature extraction and defect target detection.

[0027] Specifically, as shown in Figure 2 the second defect detection model comprises a dynamic routing module, a double-channel edge increase module, a semantic feature extraction module and a detection head module connected in sequence; the double-channel edge increase module comprises a first texture edge enhancement branch and a second texture edge enhancement branch connected in parallel, and the first texture edge enhancement branch is used for texture edge enhancement of high-texture images, and the second texture edge enhancement branch is used for texture edge enhancement of high-texture images.

[0028] The dynamic routing module is used for calculating the gray information entropy based on the proportion of the number of pixels of different gray levels in the static differential image, and inputting the static differential image into the first texture edge enhancement branch or the second texture edge enhancement branch in the double-channel edge increase module according to the size relationship between the gray information entropy and a preset information entropy threshold. In actual application, the dynamic routing module inputs the static differential image into the first texture edge enhancement branch when the gray information entropy is greater than the preset information entropy threshold, and inputs the static differential image into the second texture edge enhancement branch when the gray information entropy is less than the preset information entropy threshold. As input, the gray information entropy of the static difference image is introduced to measure the texture complexity of the image to determine which way to use for texture edge enhancement. Considering that the higher the gray information entropy, the more dispersed the image gray distribution and the more intense the change, which usually corresponds to the characteristics of the region with richer texture or more complex structure, the embodiment preferably determines that the gray information entropy is greater than the preset information entropy threshold, and then sends the static difference image to the first texture edge enhancement branch in the double-channel edge increase module for processing, otherwise, the static difference image is sent to the second texture edge enhancement branch in the double-channel edge increase module for processing. It should be noted that the preset information entropy threshold can be valued according to the actual application requirement, which is not limited here; the gray information entropy of the static difference image is calculated as follows: wherein, is the proportion of the number of pixels with the gray level of in the static difference image ; and is the gray information entropy of the static difference image .

[0029] The double-channel edge increase module is used for adaptive texture edge enhancement of the static difference image based on the texture complexity of the static difference image to obtain the corresponding texture enhancement feature; wherein, the first texture edge enhancement branch as shown in Figure 3 includes the sequentially connected hole convolution layer, sub-pixel convolution layer and channel attention layer, and the hole convolution layer uses the hole convolution with the expansion rate of 3 to perform feature extraction on , so as to capture the context information at a long distance and avoid increasing too many parameter quantities, and the obtained intermediate feature is , so as to expand the receptive field on the premise of complete detail information; in actual application, the hole convolution layer sequentially performs hole convolution processing, batch normalization (BN) processing and ReLU activation processing on the static difference image to obtain the feature , which can be expressed as: wherein, is the dilated convolution with the expansion rate of 3; is the batch normalization processing function; is the activation function.

[0030] ​Subpixel convolutional layers employ efficient subpixel convolution (ESPC) to achieve a finer resolution enhancement through channel expansion and subpixel rearrangement, significantly improving the ability to restore potential high-frequency details, i.e., enhancing high-frequency texture information, and further improving the model's ability to represent fine-grained defect features. In practical applications, subpixel convolutional layers sequentially process features... Features are obtained by performing multi-channel convolution processing, batch normalization processing, sub-pixel convolution processing, and the first fully connected layer. The corresponding calculation formula is as follows: in, This process involves reorganizing the channels in a feature map and sequentially reassembling them into a high-resolution image. Multichannel convolution; This is the learnable parameter matrix corresponding to the first fully connected layer.

[0031] The channel attention layer introduces a channel attention mechanism to assign higher weights to important texture channels in order to obtain response features that highlight micro-defect regions. In practical applications, the channel attention layer focuses on features. Global average pooling and second fully connected layer processing are performed sequentially. Activation processing, third fully connected processing (third fully connected layer) and Activation process yields features , can be represented as: in, and For activation functions; and These are the learnable parameter matrices for the second and third fully connected layers, respectively; ⊙ represents the element-wise multiplication operation; GAP This is global average pooling.

[0032] In this embodiment, the first texture edge enhancement branch can enhance fine-grained textures such as cracks and scratches in components such as insulators based on sub-pixel convolution and channel attention mechanisms, which can improve the reliability of defect target detection in high-texture images.

[0033] The second texture edge enhancement branch is as follows Figure 4 As shown, it includes a depthwise convolutional layer, an attention layer, and a global pooling layer connected in sequence. The depthwise convolutional layer uses lightweight depthwise separable convolution to extract the main structural changes in the static difference image, reduce noise interference, and obtain features. In practical applications, the deep convolutional layer is applied to the static difference image The deep separable convolution processing, batch normalization processing and ReLU activation processing are sequentially performed to obtain the feature , which is expressed as: wherein, is the deep separable convolution, and 5*5 convolution is independently performed on each channel to emphasize the spatial features of the channel.

[0034] The attention layer includes an improved window attention layer and an improved shifted window self-attention layer connected in sequence, so as to effectively capture the long-distance edge dependence and structural relationship in the feature map In the embodiment, the improved window attention layer is obtained by improving the existing lightweight Swin-Transformer window attention mechanism (Window Self-Attention, W-MSA), and the improved shifted window self-attention layer is obtained by improving the existing shifted window self-attention mechanism (Shifted Window Self-Attention, SW-MSA), and in order to obtain a higher response in the real edge area, the attention weight in the window attention mechanism and the shifted window self-attention mechanism is modulated based on the gradient of the feature map. In practical applications, the window attention layer divides the input feature into fixed-size non-overlapping windows, fuses the attention scores independently calculated in each window based on the improved window attention mechanism and the feature , and then sequentially performs layer normalization and fourth full connection processing (fourth full connection layer) to capture the local edge features , and then the feature is fused with the feature based on the improved shifted window self-attention mechanism to realize information exchange between adjacent windows, and the corresponding self-attention scores are obtained, and then the feature is fused with the feature to obtain the feature

[0035] that can highlight the edge response area. wherein, is the gradient of the feature generated by the improved window attention mechanism; The calculation formula of the attention coefficient in the corresponding window is: wherein, and are gradient weight coefficient and scaling coefficient in the improved window attention mechanism respectively; is a parameter matrix for calculating the attention coefficient; is the attention coefficient matrix of generated based on the improved window attention mechanism.

[0036] The gradient calculation formula of the feature map in each window in the improved shift window mechanism is as follows: wherein, is the gradient of generated based on the improved shift window attention mechanism; The calculation formula of the attention coefficient in the corresponding window is as follows: wherein, and are gradient weight coefficient and scaling coefficient in the improved shift window mechanism respectively; is a parameter matrix for calculating the attention coefficient; is the attention coefficient matrix of generated based on the improved shift window self-attention mechanism.

[0037] The global pooling layer performs global average pooling processing, sixth full connection processing (sixth full connection layer) and activation processing on the feature in turn to suppress invalid information, and obtains enhanced edge feature , which is expressed as follows: wherein, and are the learnable weight parameters of the fourth full connection layer, the fifth full connection layer and the sixth full connection layer respectively; ( ) is layer normalization; is global average pooling; is an activation function.

[0038] The second texture edge enhancement branch in the embodiment mainly aims at low-texture area images such as wires, and adopts the improved window attention mechanism and the shift window self-attention mechanism to strengthen the structural shape features and edge contour changes, so as to facilitate the reliability of defect target detection of low-texture images.

[0039] The semantic feature extraction module is configured to perform multi-scale semantic feature extraction on the texture-enhanced features to obtain corresponding multi-scale semantic features. Or The multi-scale semantic feature extraction is performed, and the specific extraction process can be implemented by referring to related prior art, which is not described in detail here.

[0040] The detection head module is configured to perform global average pooling and full connection processing on the multi-scale semantic features in sequence to obtain the static detection result.

[0041] The second defect detection model with the above structure realizes fine-grained edge enhancement based on texture distribution difference through the double-channel edge adaptive enhancement mechanism guided by the difference image, which can effectively ensure the precision and robustness of the defect detection model in multi-type defect detection. It should be noted that in actual application, the construction process of the second defect detection model can be understood as follows: obtaining abnormal visual images and normal visual images of different defect targets collected by a static perception device, and labeling to obtain an image dataset; then, the difference images of each image in the image dataset are obtained by using the aforementioned static difference image acquisition method, and a training set is generated; finally, based on the obtained training set, the initial network model with the sequentially connected dynamic routing module, double-channel edge increasing module, semantic feature extraction module and detection head module is trained and optimized by using the existing network model training method until the corresponding training termination condition is reached, that is, the second defect detection model with stable model parameters can be obtained.

[0042] The embodiment not only can generate a static difference image through a static difference image generation mechanism based on weighted fusion of a pixel difference image and a structure difference image, from the aspect of obtaining a defect detection image that effectively resists external light interference and highlights structural defects, to improve the robustness of local defect detection to light changes, but also can perform adaptive fine enhancement of the texture edge of the static difference image based on the texture complexity difference texture edge enhancement mechanism, from the aspect of strengthening edge texture features and highlighting structural defects, to improve the accuracy of local defect detection and recognition, and thus effectively improve the detection precision and robustness of local multi-type defects.

[0043] S12, when the static detection result meets the preset detection accuracy requirement, taking the static detection result as the target inspection result; wherein, the preset detection accuracy requirement can be determined according to actual application requirements, for example, the confidence score of the detection result can be set to exceed the corresponding confidence threshold, which can also be dynamically adjusted according to actual conditions. That is, in actual application, local defect detection analysis is preferentially based on the static visual image collected by the static perception device, and when the obtained static detection result meets the preset detection accuracy requirement, it is directly used as the final target inspection result, without occupying communication resources to upload to the cloud for detection analysis. Only when local detection cannot give accurate recognition results, the static visual image collected by the static perception device will be uploaded to the cloud for deep defect detection analysis. This can ensure the effect of power line inspection while effectively improving the efficiency of defect detection during inspection, and effectively saving cloud computing resources and communication bandwidth resources.

[0044] S13, in the case where the static detection result does not meet the preset detection accuracy requirement, the static visual image is compressed and uploaded to the cloud, and the unmanned aerial vehicle is triggered to perform image supplementing task, and the dynamic inspection image data set collected according to the image supplementing task is compressed and uploaded to the cloud, so that the cloud performs defect recognition analysis based on the first defect detection model pre-constructed according to the static visual image and the dynamic inspection image data set, to obtain the target inspection result; the first defect detection model is used for multi-view feature fusion detection of input image; wherein, the input image can be understood as each frame image in the static visual image and the dynamic inspection image data set received by the cloud after compression processing.

[0045] In actual application, when it is determined that the static detection result does not meet the preset detection accuracy requirement, it is considered that the unmanned aerial vehicle inspection track path capable of realizing comprehensive detection of the target component needs to be generated according to the deployment position of the static perception device collecting the static visual image, and the unmanned aerial vehicle equipped with high-definition camera and multi-modal sensor (such as infrared, thermal imaging, etc.) is started based on the unmanned aerial vehicle inspection track path to perform image supplementing, so as to effectively cover the line blind area that cannot be photographed by the static perception device, and the static visual image and the dynamic inspection image data set (including multiple image frames) collected by the unmanned aerial vehicle are subjected to local compression processing before being uploaded to the cloud for deep detection analysis, so as to ensure the reliability of the inspection defect detection result. It should be noted that the compression processing includes compression (such as JPEG encoding, resolution downsampling, etc.), size normalization and format standardization, etc. The size of the image formed after compression processing is , and ​standardized image data of height, width and channel number of the image respectively, to reduce the communication load and delay, improve the overall data transmission efficiency, and ensure the uniformity of the image data format processed by the cloud, and improve the efficiency of the cloud image analysis and processing.

[0046] The first defect detection model can be understood as a network model for analyzing and processing input images to identify and locate defect regions. Considering that existing defect detection using only a single view image cannot fully exploit the complementarity and semantic association between multi-view and multi-temporal information, resulting in insufficient defect target recognition accuracy in complex structures, variable poses, or partially occluded scenes, and prone to false positives and false negatives, in order to fully exploit semantic features in power line images and utilize the interaction between different view features to improve the applicability and robustness of defect detection in complex power grid environments, the embodiment preferably uses a first defect detection model based on dynamic graph convolution network enhancement and multi-channel attention fusion to perform multi-view collaborative analysis to capture more comprehensive and rich defect features, effectively improving the comprehensiveness and accuracy of defect recognition.

[0047] Specifically, the first defect detection model, as shown in Figure 5 includes a multi-view image conversion module, a multi-view feature extraction module, a feature enhancement and fusion module, and a defect target detection module connected in sequence. The multi-view image conversion module can be understood as a processing module that considers that the vertical overhead view can fully present the line direction and overall layout of the tower, which is beneficial for global structure analysis, the diagonal oblique view can take into account the height and horizontal information, which can enhance the understanding of three-dimensional spatial structure, and the side view can highlight the longitudinal arrangement and height features of the tower and conductor. In order to fully exploit semantic information in model input images and provide multi-view spatial structure information for subsequent processing, the module is used to convert multiple view images of the input image to obtain a corresponding multi-view image set. Each input image is input into three parallel view conversion branches for processing to obtain a multi-view image set including a vertical overhead view image, a diagonal oblique view image, and a side view image, which can be represented as: In the formula, wherein, is the pth input image of the multi-view image conversion module; , and are the vertical overhead view image, the diagonal oblique view image, and the side view image corresponding to the pth input image, respectively. is an affine transformation matrix for generating a vertical top-down perspective image, , is a scaling coefficient, , is a translation parameter for adjusting the overall scale and position of the image, enhancing the perception of the overall structure of the image (such as the direction of the conductor); is a perspective transformation matrix for generating a diagonal oblique perspective image, , , and are rotation and scaling components, , are translation amounts, , are perspective projection parameters for realizing the perspective deformation of the viewing angle, enhancing the modeling ability of the model for spatial inclined structures; is an affine transformation matrix for generating a side view perspective image for observing the tower and conductor, is a horizontal direction shear coefficient, is a vertical scaling coefficient to highlight the height and edge profile features of the object. It should be noted that the coefficients in the above affine transformation matrix and perspective transformation matrix can be embedded as learnable variables in the model structure. In the initial model, they are all set to 0.5, and then updated and optimized during the model training process.

[0048] The multi-view feature extraction module is configured to extract features from the vertical top-down perspective image based on a dilated spatial pyramid network, extract features from the diagonal oblique perspective image based on a deformable convolution network, and extract features from the side view perspective image based on vertical stripe pooling combined with horizontal direction convolution, to obtain corresponding multi-view feature maps; the multi-view feature maps include a vertical top-down perspective feature map, a diagonal oblique perspective feature map, and a side view perspective feature map, and all the view feature maps are size intermediate feature tensors; in actual application, the process of obtaining multi-view feature maps is as follows: 1) For the vertical top-down perspective image, since its geometric structure is complete, a dilated spatial pyramid network is used to model the structure of different scale targets based on multiple dilated convolutions with different receptive field sizes, to obtain a vertical top-down perspective feature map with stable spatial representation ability, and the specific calculation formula is as follows: wherein, is a set composed of n dilated rates; denotes a dilated convolution operation with a dilated rate of r; is a vertical top-down perspective feature map corresponding to the pth input image.

[0049] 2) For the diagonal strabismus view image, a deformable convolution network is adopted to capture the geometric deformation under the diagonal strabismus view by adaptive offset sampling points, to enhance the robustness to view changes, to obtain the diagonal strabismus view feature map, and the calculation formula is as follows: wherein, is the weight of the th convolution kernel; is the function of sampling the pixel at a certain position of the image; is the center point coordinate corresponding to the current output position; is the standard offset of the th sampling point in the convolution kernel, is the learnable offset; is the total number of convolution kernels; is the diagonal strabismus view feature map corresponding to the th input image. 3) In the side view image, the tower body and the auxiliary components are often arranged along the longitudinal direction, and the defect information (such as broken strands, loose strands, and foreign objects hanging) has obvious context association in the longitudinal direction; in order to effectively utilize this directional feature, vertical stripe pooling is adopted to compress the feature space in the vertical direction, to enhance the modeling ability of the model to the target height feature, and to cooperate with the horizontal direction convolution processing to strengthen the outline and height feature of the object, to obtain the side view feature map, and the calculation formula is as follows:

[0050] wherein, is the vertical stripe pooling operation; is the horizontal direction convolution; represents channel splicing; is the convolution kernel size; is the side view feature map corresponding to the th input image. The embodiment considers the feature difference of different view images, and selectively selects different view feature map extraction methods, which can effectively ensure the reliability of different view feature extraction.

[0051] The embodiment considers the feature difference of different view images, and selectively selects different view feature map extraction methods, which can effectively ensure the reliability of different view feature extraction.

[0052] ​​The feature enhancement fusion module can be understood as a processing module for feature enhancement and fusion of vertical overhead view feature maps, diagonal oblique view feature maps, side view feature maps and corresponding input images, for effectively mining the interaction relationship between multi-view features. The feature enhancement fusion module is used for feature enhancement and fusion of the multi-view feature maps and the input images based on a preset graph attention network and a multi-channel attention mechanism, to obtain corresponding multi-view fusion features. In order to enhance effective feature interaction and suppress information redundancy, the embodiment preferably first models the spatial correlation between different views through a graph attention network, and then realizes adaptive weighted fusion of different features by using a multi-channel attention mechanism based on tensor modal product.

[0053] Specifically, the feature enhancement fusion module includes a graph convolution processing unit, a channel splicing unit, an attention fusion unit and a boundary enhancement unit connected in sequence. The graph convolution processing unit is used to construct a full connection graph with the multi-view feature maps and the input images as nodes, and to update the multi-view feature maps and the input images according to the adjacency matrix of the full connection graph and the preset graph attention network, to obtain updated multi-view feature maps and updated input images. The construction of the full connection graph can be realized by referring to existing graph construction technology, which is not described in detail here. The edge weight of the full connection graph is adaptively learned by the graph attention network, and the update calculation process of the multi-view feature maps and the input images (which can be understood as the feature maps of the original collection view) is as follows: wherein, is the edge weight between nodes and node ; is an activation function; and are learnable weight parameter matrices; and are feature maps of node and node , respectively; is the adjacency matrix of the full connection graph, and the weight parameter between nodes is ; is one of the adjacent nodes in the adjacent node set of node ; is a graph attention network; are the updated vertical overhead view feature maps, diagonal oblique view feature maps, side view feature maps and input images corresponding to the pth input image, respectively.

[0054] The channel splicing unit is configured to splice the updated multi-view feature map and the updated input image in the channel to obtain a corresponding joint feature tensor, where the joint feature tensor is represented as: wherein, is channel splicing, is a convolutional layer; is the joint feature tensor corresponding to the pth input image.

[0055] The attention fusion unit can be understood as a processing module for adaptively assigning attention weight coefficients of different view features, which calculates the attention weight coefficients of each feature in the joint feature tensor based on a multi-channel attention mechanism of tensor modal product, and fuses the updated multi-view feature map and the updated input image based on the attention weight coefficients of each feature to obtain a corresponding enhanced feature tensor, where the calculation process of the enhanced feature tensor is as follows: 1) The multi-channel attention mechanism based on tensor modal product adaptively captures the weight occupied by each feature in the joint feature tensor in each dimension, and the attention weight coefficient of each feature is: wherein, is an activation function, is a learnable parameter matrix based on tensor modal product used in attention coefficient calculation, represents the operation of multiplying the tensor and the learnable parameter matrix along the pth dimension of the tensor; are the attention weight coefficients of the updated vertical overhead view feature map, the diagonal oblique view feature map, the side view feature map, and the input image, respectively.

[0056] 2) The attention weight coefficients are multiplied by the corresponding features and added to obtain the fused enhanced feature tensor: wherein, is a corresponding element multiplication operation; is the enhanced feature tensor corresponding to the pth input image.

[0057] The boundary enhancement unit can be understood as a double-pooling channel boundary enhancement processing module for further improving feature boundary information, which is configured to perform average pooling and maximum pooling on the enhanced feature tensor respectively, and then perform convolution processing on the first and second pooled features obtained correspondingly, and fuse the enhanced feature tensor to obtain the multi-view fusion feature. That is, the enhanced feature tensor ​​The average pooling and the maximum pooling are respectively performed, the first pooling feature is obtained by retaining more background information through the average pooling, the second pooling feature is obtained by retaining more texture information through the maximum pooling, and then the convolution operation is respectively performed on the first pooling feature and the second pooling feature to generate two branch features corresponding to the first pooling feature and the second pooling feature and After that, in order to prevent information loss, the two branch features are added and then multiplied with the original enhanced feature tensor to obtain the final multi-view fusion feature : In the formula, and are the maximum pooling and the average pooling, respectively.

[0058] The defect target detection module is configured to perform target defect recognition based on a preset neural network according to the multi-view fusion feature, to obtain a corresponding defect detection result; wherein the preset neural network can select YOLOv8 as a backbone structure to complete the boundary box prediction and the class classification of the target defect region, to obtain the defect detection result including the position coordinates and the confidence score of the defect target boundary box.

[0059] The first defect detection model with the above structure converts a single-view image into a multi-view image based on affine transformation and perspective transformation, extracts different view features, fully excavates the interaction relationship between different view features based on a graph attention network and a multi-channel attention mechanism, and performs multi-view feature fusion analysis mechanism of deep fusion and semantic enhancement on the multi-view features. Compared with the existing single-view recognition mode, the first defect detection model can more comprehensively capture defect features under complex structures or local occlusions with changing postures, and effectively improves the accuracy and comprehensiveness of defect recognition and detection in complex inspection scenes. It should be noted that in actual application, the construction process of the first defect detection model can be understood as follows: obtaining a model training set labeled by abnormal line images and normal line images of different defect targets; based on the model training set, using an existing network model training method to train and optimize the network model with the multi-view image conversion module, the multi-view feature extraction module, the feature enhancement fusion module and the defect target detection module connected in sequence, until the corresponding training termination condition is reached, that is, the first defect detection model with stable model parameters can be obtained.

[0060] In practical applications, multi-view fusion detection analysis of each frame image in the static visual image and dynamic inspection image dataset based on the first defect detection model can obtain a defect detection result in the cloud. In order to ensure that the cloud detection ultimately obtains a reliable target inspection result, the embodiment preferably comprehensively analyzes the defect detection results corresponding to each frame image in the static visual image and dynamic inspection image dataset to determine the target inspection result. Specifically, the cloud performs defect detection based on the first defect detection model pre-constructed according to the static visual image and the dynamic inspection image dataset, and obtains the target inspection result, which includes the following steps: Based on the first defect detection model, defect detection analysis is performed on each frame image in the static visual image and the dynamic inspection image dataset, and a plurality of defect detection results are obtained. The number of defect detection results corresponds to the total number of images involved in the static visual image and the dynamic inspection image dataset. The acquisition process of each defect detection result can refer to the data processing process of each functional module in the first defect detection model, which will not be described here.

[0061] When all the defect detection results are non-defect targets, the target inspection result is set to normal inspection. That is, when all the defect detection results are normal, it is considered that the inspection is normal.

[0062] When there are defect targets in all the defect detection results, the defect detection result with the highest confidence is obtained as the target inspection result. That is, if there is at least one defect detection result that is a defect target, the defect detection result with the highest confidence is taken as the final target inspection result.

[0063] The embodiment of the application provides a technical scheme for obtaining a static visual image of a target power transmission line, generating a corresponding static differential image based on the static visual image, performing local defect detection according to the static differential image to obtain a corresponding static detection result, taking the static detection result as a target inspection result when the static detection result meets a preset detection precision requirement, uploading the static visual image to the cloud when the static detection result does not meet the preset detection precision requirement, triggering a UAV to perform an image supplement task, uploading dynamic inspection image data sets collected according to the image supplement task to the cloud, and enabling the cloud to perform defect recognition analysis based on a first defect detection model for multi-view feature fusion detection of input images to obtain the target inspection result. Based on the local efficient defect detection achieved by deploying a static sensing device, the multi-view intelligent sensing architecture of dynamic and static combination and edge-cloud multi-source collaboration of the UAV emergency image supplement combined with the cloud multi-view fusion detection analysis can not only improve the spatial coverage capability and time continuity of line inspection, but also effectively improve the efficiency and accuracy of line inspection, thereby improving the robustness and real-time response capability of intelligent inspection, and meeting the intelligent inspection requirements in complex power transmission environments.

[0064] It should be noted that although each step in the above flowchart is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps has no strict order limitation, and these steps can be executed in other orders.

[0065] In one embodiment, as shown in Figure 6 A power transmission line intelligent inspection system based on UAV and cloud edge collaboration is provided, and the system comprises: A local detection module 1 is configured to obtain a static visual image of a target power transmission line, generate a corresponding static differential image based on the static visual image, and perform local defect detection according to the static differential image to obtain a corresponding static detection result. An accuracy analysis module 2 is configured to take the static detection result as a target inspection result when the static detection result meets a preset detection precision requirement. The cloud detection module 3 is configured to, in a case where the static detection result does not meet the preset detection precision requirement, compress and upload the static visual image to the cloud, trigger the UAV to perform an image supplement collection task, and compress and upload a dynamic inspection image data set collected according to the image supplement collection task to the cloud, so that the cloud performs defect identification analysis based on a first defect detection model pre-constructed based on the static visual image and the dynamic inspection image data set, to obtain the target inspection result; and the first defect detection model is configured to perform multi-view feature fusion detection on an input image.

[0066] The specific limitations of the power transmission line intelligent inspection system based on the cooperation of the UAV and the cloud edge can be seen in the limitations of the power transmission line intelligent inspection method based on the cooperation of the UAV and the cloud edge, and the corresponding technical effects can also be obtained equally, which will not be repeated here. The various modules in the power transmission line intelligent inspection system based on the cooperation of the UAV and the cloud edge can be realized by software, hardware, and combinations thereof, in whole or in part. The above-mentioned various modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the above-mentioned various modules.

[0067] In summary, the power transmission line intelligent inspection method and system based on the cooperation of the UAV and the cloud edge provided by the embodiments of the present application, on the basis of realizing local efficient defect detection based on the deployment of static sensing devices, adopts a dynamic and static combination and edge-cloud multi-source cooperation multi-view intelligent sensing architecture of emergency image supplement collection by the UAV combined with cloud multi-view fusion detection analysis, which not only improves the spatial coverage capability and time continuity of line inspection, but also effectively improves the efficiency and accuracy of line inspection, thereby improving the robustness and real-time response capability of intelligent inspection, and meeting the intelligent inspection requirements in complex power transmission environments.

[0068] Each of the embodiments in the specification is described in a progressive manner, and the directly same or similar parts of each embodiment can be referred to each other, and each embodiment mainly describes the difference from other embodiments. Especially, for the system embodiment, since it is basically similar to the method embodiment, it is described more simply, and the related parts can be referred to the part of the description of the method embodiment. It should be noted that, each technical feature of the above-mentioned embodiments can be combined arbitrarily, in order to make the description simple, not all possible combinations of the technical features of the above-mentioned embodiments are described, however, as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the description.

[0069] The above embodiments only express several preferred embodiments of the present application, which are described in more detail and in more detail, but cannot be understood as limiting the scope of the patent. It should be noted that for ordinary skilled in the art, several improvements and replacements can be made without departing from the technical principles of the present application, and these improvements and replacements should also be considered as the protection scope of the present application. Therefore, the protection scope of the present application patent should be subject to the protection scope of the claims.

Claims

1. A method for intelligent inspection of power transmission lines based on unmanned aerial vehicles (UAVs) and cloud-edge collaboration, characterized in that, The method includes: Acquire a static visual image of the target transmission line, generate a corresponding static differential image based on the static visual image, and perform local defect detection based on the static differential image to obtain the corresponding static detection result; When the static detection result meets the preset detection accuracy requirement, the static detection result is taken as the target inspection result; If the static detection result does not meet the preset detection accuracy requirement, the static visual image is compressed and uploaded to the cloud, and the UAV is triggered to perform an image re-acquisition task. The dynamic inspection image dataset collected according to the image re-acquisition task is compressed and uploaded to the cloud, so that the cloud can perform defect identification analysis based on the static visual image and the dynamic inspection image dataset, based on the pre-built first defect detection model, to obtain the target inspection result; the first defect detection model is used to perform multi-view feature fusion detection on the input image.

2. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 1, characterized in that, The step of generating a corresponding static difference image based on the static visual image includes: The static visual image is preprocessed to obtain a preprocessed static image; Based on the lighting conditions of the preprocessed static image, a corresponding static viewpoint reference image is obtained from a preset normal static visual image library. Pixel difference image and structure difference image are generated based on the preprocessed static image and the static viewpoint reference image; The pixel difference image and the structural difference image are weighted and fused to obtain the static difference image.

3. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 1, characterized in that, The step of performing local defect detection based on the static difference image to obtain the corresponding static detection result includes: The static difference image is input into a pre-constructed second defect detection model for defect identification and analysis to obtain the static detection result. The second defect detection model is used to perform adaptive texture edge enhancement on the static difference image based on the texture complexity of the static difference image, and then perform semantic feature extraction and defect target detection in sequence.

4. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 3, characterized in that, The second defect detection model includes a dynamic routing module, a dual-channel edge enhancement module, a semantic feature extraction module, and a detection head module connected in sequence; the dual-channel edge enhancement module includes a first texture edge enhancement branch and a second texture edge enhancement branch in parallel. The dynamic routing module is used to calculate grayscale information entropy based on the proportion of pixels at different grayscale levels in the static difference image, and input the static difference image into the first texture edge enhancement branch or the second texture edge enhancement branch in the dual-channel edge enhancement module according to the relationship between the grayscale information entropy and the preset information entropy threshold. The dual-channel edge enhancement module is used to adaptively enhance the texture edges of the static difference image based on the texture complexity of the static difference image, and obtain the corresponding texture enhancement features. The semantic feature extraction module is used to extract multi-scale semantic features from the texture enhancement features to obtain the corresponding multi-scale semantic features. The detection head module is used to obtain the static detection result by sequentially performing global average pooling and fully connected processing on the multi-scale semantic features.

5. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 4, characterized in that, The first texture edge enhancement branch includes a dilated convolutional layer, a subpixel convolutional layer, and a channel attention layer connected in sequence; The second texture edge enhancement branch includes a depth convolutional layer, an attention layer, and a global pooling layer connected in sequence; the attention layer includes an improved window attention layer and an improved shifted window self-attention layer connected in sequence.

6. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 5, characterized in that, Both the improved window attention layer and the improved shift window self-attention layer modulate the attention weights based on the gradient of the feature map.

7. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 1, characterized in that, The first defect detection model includes a multi-view image conversion module, a multi-view feature extraction module, a feature enhancement and fusion module, and a defect target detection module connected in sequence. The multi-view image conversion module is used to perform multi-view image conversion on the input image to obtain a corresponding multi-view image set; the multi-view image set includes vertical top view image, diagonal oblique view image and side view image; The multi-view feature extraction module is used to extract features from the vertical top-view image based on a hollow spatial pyramid network, to extract features from the diagonal oblique-view image based on a deformable convolutional network, and to extract features from the side-view image based on vertical stripe pooling combined with horizontal convolution, to obtain corresponding multi-view feature maps; the multi-view feature maps include a vertical top-view feature map, a diagonal oblique-view feature map, and a side-view feature map. The feature enhancement and fusion module is used to perform feature enhancement and fusion on the multi-view feature map and the input image based on a preset graph attention network and a multi-channel attention mechanism to obtain the corresponding multi-view fused features. The defect target detection module is used to identify target defects based on the multi-view fusion features and a preset neural network to obtain the corresponding defect detection results.

8. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 7, characterized in that, The feature enhancement and fusion module includes a graph convolution processing unit, a channel splicing unit, an attention fusion unit, and a boundary enhancement unit connected in sequence. The graph convolution processing unit is used to construct a fully connected graph using the multi-view feature map and the input image as nodes, and to update the multi-view feature map and the input image according to the adjacency matrix of the fully connected graph and the preset graph attention network to obtain the updated multi-view feature map and the updated input image. The channel stitching unit is used to stitch the updated multi-view feature map and the updated input image together to obtain the corresponding joint feature tensor. The attention fusion unit is used to calculate the attention weight coefficients of each feature in the joint feature tensor based on the multi-channel attention mechanism of tensor modality product, and to fuse the updated multi-view feature map and the updated input image based on the attention weight coefficients of each feature to obtain the corresponding enhanced feature tensor. The boundary enhancement unit is used to perform average pooling and max pooling on the enhancement feature tensor, and then convolve the corresponding first pooling feature and second pooling feature, and then fuse them with the enhancement feature tensor to obtain the multi-view fused feature.

9. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 1, characterized in that, The steps by which the cloud platform performs defect detection based on the static visual image and the dynamic inspection image dataset, using a pre-built first defect detection model, to obtain the target inspection result include: Based on the first defect detection model, defect detection analysis is performed on each frame image in the static visual image and the dynamic inspection image dataset to obtain several defect detection results. When all the defect detection results indicate that there is no defective target, the target inspection result is set to normal. When a defective target is found among all the defect detection results, the defect detection result with the highest confidence level is taken as the target inspection result.

10. A smart inspection system for power transmission lines based on unmanned aerial vehicles (UAVs) and cloud-edge collaboration, characterized in that, The system includes: The local detection module is used to acquire static visual images of the target transmission line, generate corresponding static differential images based on the static visual images, and perform local defect detection based on the static differential images to obtain corresponding static detection results. The accuracy analysis module is used to take the static detection result as the target inspection result when the static detection result meets the preset detection accuracy requirements. The cloud-based detection module is used to compress and upload the static visual image to the cloud when the static detection result does not meet the preset detection accuracy requirement, and to trigger the UAV to perform an image re-acquisition task. The module also compresses and uploads the dynamic inspection image dataset collected according to the image re-acquisition task to the cloud, so that the cloud can perform defect identification analysis based on the static visual image and the dynamic inspection image dataset, using a pre-built first defect detection model, to obtain the target inspection result. The first defect detection model is used to perform multi-view feature fusion detection on the input image.

Citation Information

Patent Citations

  • Dynamic and static cooperative power transmission line refined inspection method and system

    CN115220479A

  • Power transmission line body defect detection method based on cloud edge cooperation

    CN117975305A

  • Power transmission line key component defect identification method based on cloud edge cooperation

    CN120470463A

  • Power equipment defect detection system and method based on deep learning

    CN120526328A

  • Circuit board defect analysis method based on visual inspection

    CN120594532A

Cited By

  • Power transmission line hidden danger identification method based on wire width correction and differential shooting

    CN121767761A

  • A transmission line hidden danger identification method based on wire width correction and differential shooting

    CN121767761B

  • Visual detection method and system for repairing broken lines of airborne electric wires

    CN121883492A