Power transmission line intelligent inspection method and system based on unmanned aerial vehicle and cloud edge cooperation
By using a smart inspection method that combines drones and cloud-edge collaboration, along with static sensing equipment and cloud-based multi-view analysis, the problems of fixed viewing angles and detection delays in power transmission line inspections have been solved. This has enabled efficient and accurate defect detection, improving the coverage and real-time performance of inspections.
Patent Information
- Application Number
- CN202511494858.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-10-20
AI Technical Summary
Existing transmission line inspection methods suffer from problems such as fixed viewing angles, insufficient coverage, detection delays, and insufficient defect identification accuracy. In particular, they are difficult to achieve high-frequency, continuous inspections and real-time detection in complex environments.
The intelligent inspection method adopts drones and cloud-edge collaboration. Local defect detection is carried out through static sensing devices, and the image acquisition by drones and multi-view fusion analysis in the cloud are combined to realize a multi-view intelligent sensing architecture that combines dynamic and static elements and multi-source collaboration between the edge and cloud, thereby improving the spatial coverage and temporal continuity of the inspection.
This improves the efficiency, accuracy, and robustness of power transmission line inspection, ensuring real-time response and detection accuracy for line defects.
Smart Images

Figure CN120976809B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power transmission line inspection technology, and in particular to an intelligent power transmission line inspection method and system based on unmanned aerial vehicles (UAVs) and cloud-edge collaboration. Background Technology
[0002] Transmission lines, as crucial channels for large-scale, long-distance power transmission in power systems, directly impact the stability of the power grid structure and the continuity of power supply. With the large-scale integration of new energy sources, rapid growth in electricity demand, and the continuous expansion of power grid coverage areas, the distribution of transmission lines is becoming increasingly widespread, and their operating environment is becoming increasingly complex. This leads to a significant increase in the frequency of defects, easily causing power outages, equipment damage, and even regional blackouts, seriously threatening the safe operation of the power system. Therefore, conducting efficient, accurate, and sustainable condition monitoring and defect identification of transmission lines has become one of the key tasks of current power grid operation and maintenance management.
[0003] Currently, intelligent inspection of transmission lines mainly adopts fixed static sensing equipment or inspection based on dynamic equipment such as drones to address the shortcomings of manual inspection. However, although existing intelligent inspection methods have made significant progress in transmission line defect detection, they still have limitations: 1) Static sensing equipment has a fixed perspective and cannot cover the entire line. Although drones have flexible data acquisition capabilities, they are limited by endurance and scheduling complexity, making it difficult to achieve high-frequency, continuous inspection, resulting in insufficient overall coverage and continuity; 2) Uploading all images collected by static equipment or drones to the cloud for processing often introduces delays, affecting the real-time performance of detection; 3) Defect detection models often rely on single perspectives or isolated images for discrimination, failing to fully explore the complementarity and semantic correlation between multi-perspective information, and failing to consider the differences in texture distribution features of different target images, resulting in insufficient accuracy in identifying defect targets in transmission lines with complex structures and variable postures. Summary of the Invention
[0004] The purpose of this invention is to provide an intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration. On the basis of deploying static sensing equipment to achieve efficient local defect detection, it adopts a dynamic and static combined with cloud-based multi-view fusion detection and analysis, which combines UAV emergency image acquisition and cloud-edge multi-source collaboration. This not only improves the spatial coverage and temporal continuity of line inspection, but also effectively enhances the efficiency and accuracy of line inspection, thereby improving the robustness of intelligent inspection.
[0005] To achieve the above objectives, it is necessary to provide a method and system for intelligent inspection of power transmission lines based on drones and cloud-edge collaboration.
[0006] In a first aspect, embodiments of the present invention provide a method for intelligent inspection of power transmission lines based on unmanned aerial vehicles (UAVs) and cloud-edge collaboration, the method comprising:
[0007] Acquire a static visual image of the target transmission line, generate a corresponding static differential image based on the static visual image, and perform local defect detection based on the static differential image to obtain the corresponding static detection result;
[0008] When the static detection result meets the preset detection accuracy requirement, the static detection result is taken as the target inspection result;
[0009] If the static detection result does not meet the preset detection accuracy requirement, the static visual image is compressed and uploaded to the cloud, and the UAV is triggered to perform an image re-acquisition task. The dynamic inspection image dataset collected according to the image re-acquisition task is compressed and uploaded to the cloud, so that the cloud can perform defect identification analysis based on the static visual image and the dynamic inspection image dataset, based on the pre-built first defect detection model, to obtain the target inspection result; the first defect detection model is used to perform multi-view feature fusion detection on the input image.
[0010] Furthermore, the step of generating a corresponding static difference image based on the static visual image includes:
[0011] The static visual image is preprocessed to obtain a preprocessed static image;
[0012] Based on the lighting conditions of the preprocessed static image, a corresponding static viewpoint reference image is obtained from a preset normal static visual image library.
[0013] Pixel difference image and structure difference image are generated based on the preprocessed static image and the static viewpoint reference image;
[0014] The pixel difference image and the structural difference image are weighted and fused to obtain the static difference image.
[0015] Further, the step of performing local defect detection based on the static difference image to obtain the corresponding static detection result includes:
[0016] The static difference image is input into a pre-constructed second defect detection model for defect identification and analysis to obtain the static detection result. The second defect detection model is used to perform adaptive texture edge enhancement on the static difference image based on the texture complexity of the static difference image, and then perform semantic feature extraction and defect target detection in sequence.
[0017] Furthermore, the second defect detection model includes a dynamic routing module, a dual-channel edge enhancement module, a semantic feature extraction module, and a detection head module connected in sequence; the dual-channel edge enhancement module includes a first texture edge enhancement branch and a second texture edge enhancement branch in parallel.
[0018] The dynamic routing module is used to calculate grayscale information entropy based on the proportion of pixels at different grayscale levels in the static difference image, and input the static difference image into the first texture edge enhancement branch or the second texture edge enhancement branch in the dual-channel edge enhancement module according to the relationship between the grayscale information entropy and the preset information entropy threshold.
[0019] The dual-channel edge enhancement module is used to adaptively enhance the texture edges of the static difference image based on the texture complexity of the static difference image, and obtain the corresponding texture enhancement features.
[0020] The semantic feature extraction module is used to extract multi-scale semantic features from the texture enhancement features to obtain the corresponding multi-scale semantic features.
[0021] The detection head module is used to obtain the static detection result by sequentially performing global average pooling and fully connected processing on the multi-scale semantic features.
[0022] Furthermore, the first texture edge enhancement branch includes a dilated convolutional layer, a subpixel convolutional layer, and a channel attention layer connected in sequence;
[0023] The second texture edge enhancement branch includes a depth convolutional layer, an attention layer, and a global pooling layer connected in sequence; the attention layer includes an improved window attention layer and an improved shifted window self-attention layer connected in sequence.
[0024] Furthermore, both the improved window attention layer and the improved shift window self-attention layer modulate the attention weights based on the gradient of the feature map.
[0025] Furthermore, the first defect detection model includes a multi-view image conversion module, a multi-view feature extraction module, a feature enhancement and fusion module, and a defect target detection module connected in sequence;
[0026] The multi-view image conversion module is used to perform multi-view image conversion on the input image to obtain a corresponding multi-view image set; the multi-view image set includes vertical top view image, diagonal oblique view image and side view image;
[0027] The multi-view feature extraction module is used to extract features from the vertical top-view image based on a hollow spatial pyramid network, to extract features from the diagonal oblique-view image based on a deformable convolutional network, and to extract features from the side-view image based on vertical stripe pooling combined with horizontal convolution, to obtain corresponding multi-view feature maps; the multi-view feature maps include a vertical top-view feature map, a diagonal oblique-view feature map, and a side-view feature map.
[0028] The feature enhancement and fusion module is used to perform feature enhancement and fusion on the multi-view feature map and the input image based on a preset graph attention network and a multi-channel attention mechanism to obtain the corresponding multi-view fused features.
[0029] The defect target detection module is used to identify target defects based on the multi-view fusion features and a preset neural network to obtain the corresponding defect detection results.
[0030] Furthermore, the feature enhancement and fusion module includes a graph convolution processing unit, a channel splicing unit, an attention fusion unit, and a boundary enhancement unit connected in sequence;
[0031] The graph convolution processing unit is used to construct a fully connected graph using the multi-view feature map and the input image as nodes, and to update the multi-view feature map and the input image according to the adjacency matrix of the fully connected graph and the preset graph attention network to obtain the updated multi-view feature map and the updated input image.
[0032] The channel stitching unit is used to stitch the updated multi-view feature map and the updated input image together to obtain the corresponding joint feature tensor.
[0033] The attention fusion unit is used to calculate the attention weight coefficients of each feature in the joint feature tensor based on the multi-channel attention mechanism of tensor modality product, and to fuse the updated multi-view feature map and the updated input image based on the attention weight coefficients of each feature to obtain the corresponding enhanced feature tensor.
[0034] The boundary enhancement unit is used to perform average pooling and max pooling on the enhancement feature tensor, and then convolve the corresponding first pooling feature and second pooling feature, and then fuse them with the enhancement feature tensor to obtain the multi-view fused feature.
[0035] Further, the step of obtaining the target inspection result by performing defect detection based on the static visual image and the dynamic inspection image dataset using a pre-built first defect detection model in the cloud includes:
[0036] Based on the first defect detection model, defect detection analysis is performed on each frame image in the static visual image and the dynamic inspection image dataset to obtain several defect detection results.
[0037] When all the defect detection results indicate that there is no defective target, the target inspection result is set to normal.
[0038] When a defective target is found among all the defect detection results, the defect detection result with the highest confidence level is taken as the target inspection result.
[0039] Secondly, embodiments of the present invention provide an intelligent inspection system for power transmission lines based on unmanned aerial vehicles (UAVs) and cloud-edge collaboration, the system comprising:
[0040] The local detection module is used to acquire static visual images of the target transmission line, generate corresponding static differential images based on the static visual images, and perform local defect detection based on the static differential images to obtain corresponding static detection results.
[0041] The accuracy analysis module is used to take the static detection result as the target inspection result when the static detection result meets the preset detection accuracy requirements.
[0042] The cloud-based detection module is used to compress and upload the static visual image to the cloud when the static detection result does not meet the preset detection accuracy requirement, and to trigger the UAV to perform an image re-acquisition task. The module also compresses and uploads the dynamic inspection image dataset collected according to the image re-acquisition task to the cloud, so that the cloud can perform defect identification analysis based on the static visual image and the dynamic inspection image dataset, using a pre-built first defect detection model, to obtain the target inspection result. The first defect detection model is used to perform multi-view feature fusion detection on the input image.
[0043] This invention provides an intelligent inspection method and system for power transmission lines based on UAVs and cloud-edge collaboration. The method acquires static visual images of the target power transmission line, generates corresponding static differential images based on these images, performs local defect detection using the differential images to obtain corresponding static detection results, and uses these results as the target inspection result when they meet preset detection accuracy requirements. Conversely, when these results do not meet the preset accuracy requirements, the static visual images are compressed and uploaded to the cloud, triggering the UAV to perform an image re-acquisition task. The UAV then compresses and uploads the dynamic inspection image dataset acquired according to the image re-acquisition task to the cloud. This allows the cloud to perform defect identification and analysis based on the static visual images and dynamic inspection image dataset, using a pre-constructed first defect detection model for multi-view feature fusion detection of input images, to obtain the target inspection result. Compared with existing technologies, this intelligent transmission line inspection method based on UAVs and cloud-edge collaboration, on the basis of deploying static sensing equipment to achieve efficient local defect detection, adopts a dynamic and static combined with cloud-based multi-view fusion detection and analysis, and a multi-view intelligent sensing architecture that combines static and dynamic elements and edge-cloud multi-source collaboration. This not only improves the spatial coverage and temporal continuity of line inspection, but also effectively enhances the efficiency and accuracy of line inspection, thereby improving the robustness and real-time response capability of intelligent inspection. Attached Figure Description
[0044] Figure 1 This is a schematic diagram of the process of intelligent inspection of power transmission lines based on drones and cloud-edge collaboration in an embodiment of the present invention;
[0045] Figure 2 This is a schematic diagram of the structure of the second defect detection model in this embodiment of the invention;
[0046] Figure 3 yes Figure 2 A schematic diagram of the structure of the first texture edge enhancement branch;
[0047] Figure 4 yes Figure 2 A schematic diagram of the structure of the second texture edge enhancement branch;
[0048] Figure 5 This is a schematic diagram of the structure of the first defect detection model in an embodiment of the present invention;
[0049] Figure 6 This is a schematic diagram of the intelligent power transmission line inspection system based on UAV and cloud-edge collaboration in an embodiment of the present invention;
[0050] The attached figures are labeled as follows:
[0051] 1. Local detection module; 2. Accuracy analysis module; 3. Cloud detection module. Detailed Implementation
[0052] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. Obviously, the embodiments described below are only part of the embodiments of this invention and are used to illustrate the invention, but are not intended to limit the scope of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0053] In some embodiments, such as Figure 1 As shown, a method for intelligent inspection of power transmission lines based on drones and cloud-edge collaboration is provided, including the following steps:
[0054] S11. Obtain a static visual image of the target transmission line. Based on the static visual image, generate a corresponding static differential image. Perform local defect detection based on the static differential image to obtain the corresponding static detection result. The static visual image can be understood as a fixed-view image at the deployment location, acquired at a preset sampling frequency by a static sensing device deployed on key nodes of the transmission line. The static sensing device can be an existing image acquisition device (e.g., a fixedly installed camera).
[0055] Considering that static visual image acquisition is easily affected by ambient light interference, which can impact the accuracy of defect detection based on static visual images, this embodiment preferably uses the AIBOX edge computing device to acquire differential features of static visual images based on historical normal static viewpoint images as analysis data for subsequent defect detection. This improves robustness to changes in illumination. It should be noted that the AIBOX edge computing device can perform real-time online monitoring of images acquired by a fixed camera, eliminating the need to upload all images to the cloud for detection, thus reducing detection latency. Specifically, the step of generating a corresponding static differential image based on the static visual image includes:
[0056] The static visual image is preprocessed to obtain a preprocessed static image. The preprocessing may include size normalization, pixel normalization and noise removal of the static visual image to obtain a preprocessed static image with stable data quality. The specific processing procedure is implemented with reference to relevant existing technologies and will not be described in detail here.
[0057] Based on the illumination conditions of the preprocessed static image, a corresponding static viewpoint reference image is obtained from a preset normal static visual image library. Illumination conditions include brightness and contrast, which can be obtained through feature analysis of the preprocessed static image. For example, brightness can be extracted by calculating the average grayscale value of all pixels in the image, and contrast can be obtained by statistically analyzing the variance of the image's grayscale distribution. The preset normal static visual image library can be understood as a database containing images of normal transmission lines under different illumination conditions at different key points along the transmission line. It can be constructed and updated based on normal transmission line images collected by static sensing devices deployed at various key nodes along the transmission line. The static viewpoint reference image can be understood as an image obtained by matching and analyzing the preprocessed static image with normal transmission line images under the same illumination conditions in the preset normal static visual image library. This ensures that the obtained image has similar illumination conditions to the preprocessed static image, so that the subsequent static difference image reflects only the actual defects, not the illumination differences.
[0058] The actual matching analysis process for obtaining the static viewpoint reference image can be as follows: based on the preprocessed static image corresponding to the current time t... Lighting conditions, from a preset normal static visual image library Select the static viewpoint reference image whose histogram is most similar (highest histogram similarity). Preprocessing static images and each reference image in the preset normal static visual image library histogram similarity The correlation coefficient of the histogram is obtained using the following formula:
[0059]
[0060] in, For preprocessing still images In the Histogram values at each gray level; To pre-define the k-th reference image in a normal static visual image library In the Histogram values at each gray level, where N represents the total number of reference images in the preset normal static visual image library; and Preprocessed static images and reference image The histogram average.
[0061] Based on the preprocessed static image and the static viewpoint reference image, a pixel difference image and a structural difference image are generated. The pixel difference image can be understood as an image composed of the absolute values of the pixel differences at corresponding positions in the preprocessed static image and the static viewpoint reference image, directly reflecting grayscale changes between the images. The structural difference image can be understood as a difference image constructed based on the structural similarity (SSIM) index to obtain a more robust difference image, considering that the pixel difference image is still sensitive to illumination changes and easily affected by noise. In practical applications, the local SSIM values of the preprocessed static image and the static viewpoint reference image can be calculated by sliding a preset sliding window. This yields a structural similarity matrix that reflects the local structural consistency at the same positions in the preprocessed static image and the static viewpoint reference image. Then, the obtained structural similarity matrix is upsampled using a bicubic interpolation function to obtain a structural difference image spatially aligned with the pixel difference image. The specific calculation of local SSIM values and bicubic interpolation can be found in existing technologies and will not be detailed here.
[0062] The pixel difference image and the structural difference image are weighted and fused to obtain the static difference image; wherein, the static difference image can be represented as:
[0063]
[0064] in, and These are the preprocessed static image and the corresponding static view reference image at time t, respectively. for and The structural similarity matrix; It is a bicubic interpolation upsampling function; These are dynamic weighting coefficients; for The corresponding static difference image.
[0065] This embodiment obtains a static difference image by fusing the structural similarity (SSIM) index on the basis of traditional pixel difference. This enables the evaluation of image differences while effectively suppressing the influence of illumination, further enhancing the structural perception capability and providing a reliable analytical basis for subsequent defect detection and identification.
[0066] The static difference image obtained through the above method steps can be used as the basic data for defect target detection. To ensure the reliability of defect detection, this embodiment preferably performs fine-grained edge enhancement based on fully considering the differences in texture distribution features of different target images, so as to improve the accuracy and robustness of multi-type defect detection. Specifically, the step of performing local defect detection based on the static difference image to obtain the corresponding static detection result includes:
[0067] The static difference image is input into a pre-constructed second defect detection model for defect identification and analysis to obtain the static detection result. The second defect detection model can be understood as a network model that can be used to analyze the static difference image to obtain static detection results including the target bounding box position coordinates and detection confidence scores. Considering that the high-texture regions in the image contain rich fine-grained texture information, a large receptive field and fine texture extraction capabilities are required, while the low-texture regions contain more structural dependencies, requiring the ability to effectively suppress background noise and improve the expression efficiency of structural information, this embodiment preferably adopts a network structure that can adaptively enhance the texture edge of the static difference image based on the texture complexity of the static difference image, and then sequentially perform semantic feature extraction and defect target detection.
[0068] Specifically, such as Figure 2 As shown, the second defect detection model includes a dynamic routing module, a dual-channel edge enhancement module, a semantic feature extraction module, and a detection head module connected in sequence; the dual-channel edge enhancement module includes a first texture edge enhancement branch and a second texture edge enhancement branch in parallel, and the first texture edge enhancement branch is used to enhance the texture edges of the high-texture image, and the second texture edge enhancement branch is used to enhance the texture edges of the high-texture image.
[0069] The dynamic routing module calculates grayscale entropy based on the proportion of pixels at different gray levels in the static difference image, and inputs the static difference image into the first texture edge enhancement branch or the second texture edge enhancement branch of the dual-channel edge enhancement module according to the relationship between the grayscale entropy and a preset entropy threshold. In practical applications, the dynamic routing module uses the static difference image... As input, the grayscale entropy of the static difference image is used to measure the image texture complexity, determining the appropriate method for texture edge enhancement. Considering that higher grayscale entropy indicates a more dispersed and drastic grayscale distribution, typically corresponding to regions with richer textures or more complex structures, this embodiment preferably sends the static difference image to the first texture edge enhancement branch of the dual-channel edge enhancement module for processing when the grayscale entropy is determined to be greater than a preset entropy threshold; conversely, it sends the static difference image to the second texture edge enhancement branch of the dual-channel edge enhancement module for processing. It should be noted that the preset entropy threshold can be set according to actual application requirements and is not specifically limited here; the static difference image... Gray information entropy The calculation formula is as follows:
[0070]
[0071] in, For static difference images medium gray level is The percentage of pixels; For static difference images The grayscale information entropy.
[0072] The dual-channel edge enhancement module is used to adaptively enhance the texture edges of the static difference image based on the texture complexity of the static difference image, thereby obtaining corresponding texture enhancement features; wherein, the first texture edge enhancement branch is as follows: Figure 3 As shown, it includes a dilated convolutional layer, a subpixel convolutional layer, and a channel attention layer connected in sequence, and the dilated convolutional layer uses a dilated convolutional pair with a dilation rate of 3. Feature extraction is performed to capture contextual information over longer distances while avoiding excessive parameter increases. The intermediate features obtained are... To ensure the receptive field is expanded while preserving detailed information, in practical applications, dilated convolutional layers sequentially perform dilated convolution, batch normalization (BN), and ReLU activation on static difference images to obtain features. , can be represented as:
[0073]
[0074] in, This is a dilated convolution with a dilation rate of 3; This is a batch standardization function; This is the activation function.
[0075] Subpixel convolutional layers employ efficient subpixel convolution (ESPC) to achieve a finer resolution enhancement through channel expansion and subpixel rearrangement, significantly improving the ability to restore potential high-frequency details, i.e., enhancing high-frequency texture information, and further improving the model's ability to represent fine-grained defect features. In practical applications, subpixel convolutional layers sequentially process features... Features are obtained by performing multi-channel convolution processing, batch normalization processing, sub-pixel convolution processing, and the first fully connected layer. The corresponding calculation formula is as follows:
[0076]
[0077] in, This process involves reorganizing the channels in a feature map and sequentially reassembling them into a high-resolution image. Multichannel convolution; This is the learnable parameter matrix corresponding to the first fully connected layer.
[0078] The channel attention layer introduces a channel attention mechanism to assign higher weights to important texture channels in order to obtain response features that highlight micro-defect regions. In practical applications, the channel attention layer focuses on features. Global average pooling and second fully connected layer processing are performed sequentially. Activation processing, third fully connected processing (third fully connected layer) and Activation process yields features , can be represented as:
[0079]
[0080] in, and For activation functions; and These are the learnable parameter matrices for the second and third fully connected layers, respectively; ⊙ represents the element-wise multiplication operation; GAP This is global average pooling.
[0081] In this embodiment, the first texture edge enhancement branch can enhance fine-grained textures such as cracks and scratches in components such as insulators based on sub-pixel convolution and channel attention mechanisms, which can improve the reliability of defect target detection in high-texture images.
[0082] The second texture edge enhancement branch is as follows Figure 4As shown, it includes a depthwise convolutional layer, an attention layer, and a global pooling layer connected in sequence. The depthwise convolutional layer uses lightweight depthwise separable convolution to extract the main structural changes in the static difference image, reduce noise interference, and obtain features. In practical applications, deep convolutional layers are used for static difference images. Features are obtained by sequentially performing depthwise separable convolution, batch normalization, and ReLU activation. , is represented as:
[0083]
[0084] in, For depthwise separable convolution, perform a 5×5 convolution independently on each channel to emphasize the spatial features of the channel.
[0085] The attention layer comprises an improved window attention layer and an improved shifted window self-attention layer connected in sequence to effectively capture feature maps. The embodiment employs an improved window attention layer, which is an improvement upon the existing lightweight Swin-Transformer window attention mechanism (W-MSA), and an improved shifted window self-attention layer, which is an improvement upon the existing shifted window self-attention mechanism (SW-MSA). Furthermore, to achieve higher responses in real edge regions, this embodiment preferably modulates the attention weights in the window attention mechanism and the shifted window self-attention mechanism based on the gradient of the feature map. In practical applications, the window attention layer first processes the input features... Divided into fixed-size, non-overlapping windows, attention scores are calculated independently within each window based on an improved window attention mechanism, and features are then considered. The process involves fusion, followed by layer normalization and a fourth fully connected layer to capture local edge features. Then through the features An improved shift-window self-attention mechanism is used to achieve information exchange between adjacent windows. After obtaining the corresponding self-attention scores, they are then compared with features. The mixture is then fused, followed by layer normalization and a fifth fully connected layer to obtain features that highlight the edge response regions. .
[0086] In this embodiment, the feature map gradients used in both the improved window attention mechanism and the improved shift window mechanism can be calculated using the Sobel operator. In the improved window attention mechanism, the gradient calculation formula for the feature map within each window is as follows:
[0087]
[0088] in, Generated for an improved window attention mechanism The gradient;
[0089] The formula for calculating the attention coefficient within the corresponding window is:
[0090]
[0091] in, and These are the gradient weight coefficients and scaling coefficients in the improved window attention mechanism, respectively. The parameter matrix for calculating the attention coefficients; For generating based on an improved window attention mechanism Attention coefficient matrix.
[0092] In the improved shift window mechanism, the gradient calculation formula for the feature map within each window is as follows:
[0093]
[0094] in, Generated for an improved window attention mechanism The gradient;
[0095] The formula for calculating the attention coefficient within the corresponding window is:
[0096]
[0097] in, and These are the gradient weight coefficients and scaling coefficients in the improved shift window mechanism, respectively. The parameter matrix for calculating the attention coefficients; For generating based on an improved shift-window self-attention mechanism Attention coefficient matrix.
[0098] Global pooling layers, through feature... Global average pooling, sixth fully connected layer processing (sixth fully connected layer), and... Activation processing is used to suppress invalid information and obtain enhanced edge features. , means as follows:
[0099]
[0100] in, and These are the learnable weight parameters for the fourth, fifth, and sixth fully connected layers, respectively. ( ) is for layer standardization; For global average pooling; This is the activation function.
[0101] In this embodiment, the second texture edge enhancement branch is mainly aimed at low-texture areas such as wires. It adopts an improved window attention mechanism and a shifted window self-attention mechanism to enhance structural shape features and edge contour changes, which facilitates the improvement of the reliability of defect target detection in low-texture images.
[0102] The semantic feature extraction module is used to extract multi-scale semantic features from the texture enhancement features to obtain corresponding multi-scale semantic features; wherein, the multi-scale semantic features can be obtained by using the texture enhancement features output by the dual-channel edge enhancement module through the MobileNetV3 lightweight backbone network. or The feature extraction is obtained by multi-scale semantic feature extraction. The specific extraction process can be found in existing technologies and will not be detailed here.
[0103] The detection head module is used to obtain the static detection result by sequentially performing global average pooling and fully connected processing on the multi-scale semantic features.
[0104] The second defect detection model with the above structure achieves fine-grained edge enhancement based on texture distribution differences through a dual-channel edge adaptive enhancement mechanism guided by differential images. This effectively ensures the accuracy and robustness of the defect detection model in detecting multiple types of defects. It should be noted that the construction process of the second defect detection model in practical applications can be understood as follows: acquire abnormal and normal visual images of different defect targets collected by a static sensing device, and annotate them to obtain an image dataset; then, use the aforementioned static differential image acquisition method to obtain the differential images of each image in the image dataset, generating a training set; finally, based on the obtained training set, use existing network model training methods to train and optimize the initial network model, which has a dynamically connected dynamic routing module, a dual-channel edge enhancement module, a semantic feature extraction module, and a detection head module, until the corresponding training termination condition is reached, thus obtaining the second defect detection model with stable model parameters.
[0105] This embodiment not only improves the robustness of local defect detection to illumination changes by using a static differential image generation mechanism based on weighted fusion of pixel differential images and structural differential images to obtain defect detection images that effectively resist external illumination interference and highlight structural defects, but also improves the accuracy of local defect detection and recognition by using a differentiated texture edge enhancement mechanism that adaptively and finely enhances the texture edges of the static differential image based on texture complexity, thereby strengthening edge texture features and highlighting structural defects. This effectively improves the detection accuracy and robustness of local multi-type defects.
[0106] S12. When the static detection result meets the preset detection accuracy requirement, the static detection result is used as the target inspection result. The preset detection accuracy requirement can be determined according to actual application needs. For example, it can be set as the confidence score of the detection result exceeding the corresponding confidence threshold. This confidence threshold can also be dynamically adjusted according to actual conditions. That is, in practical applications, local defect detection and analysis are prioritized based on static visual images collected by the static sensing device. When the obtained static detection result meets the preset detection accuracy requirement, it is directly used as the final target inspection result without consuming communication resources to upload it to the cloud for detection and analysis. Only when local detection fails to provide accurate identification results will the static visual images collected by the static sensing device be uploaded to the cloud for in-depth defect detection and analysis. This ensures the effectiveness of power line inspection while effectively improving the efficiency of defect detection and saving cloud computing and communication bandwidth resources.
[0107] S13. If the static detection result does not meet the preset detection accuracy requirement, the static visual image is compressed and uploaded to the cloud, and the UAV is triggered to perform an image acquisition task. The dynamic inspection image dataset acquired according to the image acquisition task is compressed and uploaded to the cloud, so that the cloud can perform defect identification analysis based on the static visual image and the dynamic inspection image dataset, based on the pre-built first defect detection model, to obtain the target inspection result. The first defect detection model is used to perform multi-view feature fusion detection on the input image. The input image can be understood as the static visual image and the various frame images in the dynamic inspection image dataset that have been compressed and have a uniform size and format received by the cloud.
[0108] In practical applications, when the static inspection results fail to meet the preset detection accuracy requirements, it is deemed necessary to generate a drone inspection trajectory path based on the deployment location of the static sensing equipment that acquires static visual images. This path allows for comprehensive inspection of the target components. Based on this trajectory path, a drone equipped with a high-definition camera and multimodal sensors (such as infrared and thermal imaging) is activated to acquire additional images, effectively covering blind spots in the line that are difficult to capture from the static sensing equipment's perspective. The static visual images and the dynamic inspection image dataset (including multiple image frames) acquired by the drone are then subjected to localized compression processing before being uploaded to the cloud for in-depth detection and analysis, ensuring the reliability of the defect detection results. It should be noted that the compression processing includes compression (e.g., JPEG encoding, resolution downsampling), size normalization, and format standardization, resulting in images with uniform dimensions. ( , and Standardized image data (including image height, width, and number of channels) is used to reduce communication load and latency, improve overall data transmission efficiency, ensure the uniformity of image data format in the cloud, and enhance the efficiency of cloud image analysis and processing.
[0109] The first defect detection model can be understood as a network model used to analyze and process input images to identify and locate defect areas. Considering that existing defect detection methods that only use single-view images cannot fully exploit the complementarity and semantic relationships between multi-view and multi-temporal information, resulting in insufficient accuracy in defect target recognition in scenarios with complex structures, varied postures, or partial occlusion, and prone to false detections and missed detections, in order to fully exploit the semantic features in transmission line images and utilize the interaction relationships between features from different perspectives to improve the applicability and robustness of defect detection in complex power grid environments, this embodiment preferably uses a first defect detection model based on dynamic graph convolutional network enhancement and multi-channel attention fusion to perform multi-view collaborative analysis, so as to capture more comprehensive and richer defect features and effectively improve the comprehensiveness and accuracy of defect recognition.
[0110] Specifically, the first defect detection model is as follows: Figure 5As shown, the system includes a multi-view image conversion module, a multi-view feature extraction module, a feature enhancement and fusion module, and a defect target detection module connected in sequence. The multi-view image conversion module can be understood as a processing module that takes into account the fact that the vertical top-down view can fully present the route and the overall layout of the towers, which is beneficial for global structural analysis; the diagonal oblique view can take into account both height and lateral information, which can enhance the understanding of the three-dimensional spatial structure; and the side view can highlight the longitudinal arrangement and height features of the towers and conductors. To fully extract the semantic information in the input images of the model, this module provides multi-view spatial structural information for subsequent processing. It is used to perform multi-view image conversion on the input images to obtain the corresponding multi-view image set. Each input image is simultaneously input into three parallel view transformation branches for processing, resulting in a multi-view image set including vertical top-down view images, diagonal oblique view images, and side view images, which can be represented as:
[0111]
[0112] In the formula,
[0113]
[0114]
[0115]
[0116] in, This is the p-th input image for the multi-view image conversion module; , and These are the vertical top-view image, the diagonal oblique-view image, and the side-view image corresponding to the p-th input image, respectively. The affine transformation matrix for generating a vertical top-down view image from directly above. , This is the scaling factor. , These are translation parameters used to adjust the overall scale and position of the image, enhancing the perception of the overall structure of the image (such as the direction of the conductor); To generate the perspective transformation matrix for the diagonal oblique view image, , , and For rotation and scaling components, , The translation amount, , These are perspective projection parameters used to achieve perspective distortion of the viewpoint and enhance the model's ability to model spatial tilted structures. To generate the affine transformation matrix for observing the side-view images of the tower and conductor, The horizontal shear coefficient is... This represents the scaling factor in the vertical direction, used to emphasize the height and edge contour features of the object. It should be noted that the coefficients in the affine transformation matrix and perspective transformation matrix can be embedded as learnable variables into the model structure. In the initial model, they are all set to 0.5, and then continuously updated and optimized during model training.
[0117] The multi-view feature extraction module is used to extract features from the vertical top-down view image based on a hollow spatial pyramid network, to extract features from the diagonal oblique view image based on a deformable convolutional network, and to extract features from the side view image based on vertical stripe pooling combined with horizontal convolution, thereby obtaining corresponding multi-view feature maps. The multi-view feature maps include a vertical top-down view feature map, a diagonal oblique view feature map, and a side view feature map, and all view feature maps are... The size of the intermediate feature tensor; in practical applications, the process of obtaining multi-view feature maps is as follows:
[0118] 1) For vertical top-down view images, due to their intact geometric structure, a dilated spatial pyramid network is used. Based on multiple dilated convolutions with different receptive field sizes, structural modeling of targets at different scales is performed to obtain a vertical top-down view feature map with stable spatial representation capabilities. The specific calculation formula is as follows:
[0119]
[0120] in, It is a set consisting of n void ratios; This represents a dilated convolution operation with a dilation rate of r. This is the vertical top-down view feature map corresponding to the p-th input image.
[0121] 2) For diagonal squint view images, a deformable convolutional network is used to capture geometric deformations under diagonal squint view by adaptively offset sampling points, thereby enhancing robustness to view changes and obtaining diagonal squint view feature maps. The calculation formula is as follows:
[0122]
[0123] in, For the first The weights of each convolutional kernel; For the image A function that samples pixels at a certain location; The coordinates of the center point corresponding to the current output position; The first convolution kernel Standard offset of each sampling point The learnable offset; This represents the total number of convolution kernels; This is the diagonal oblique view feature map corresponding to the p-th input image.
[0124] 3) In side-view images, the main body and auxiliary components of the tower are often arranged along the longitudinal direction, and defect information (such as broken strands, loose strands, and foreign objects attached) has obvious contextual relevance in the longitudinal direction. To effectively utilize this directional feature, vertical stripe pooling is used to compress the feature space in the vertical direction, thereby enhancing the model's ability to model the target height features. At the same time, horizontal convolution processing is used to strengthen the contour and height features of the object to obtain the side-view feature map. The calculation formula is as follows:
[0125]
[0126] in, This is a vertical stripe pooling operation; Convolution in the horizontal direction; Indicates channel splicing; The kernel size; This is the side view feature map corresponding to the p-th input image.
[0127] This embodiment takes into account the feature differences of images from different perspectives and selects different perspective feature map extraction methods accordingly, which can effectively ensure the reliability of feature extraction from different perspectives.
[0128] The feature enhancement and fusion module can be understood as a processing module that enhances and fuses the vertical top-view feature map, diagonal oblique view feature map, side view feature map, and corresponding input image to effectively mine the interaction relationship between multi-view features. It is used to enhance and fuse the multi-view feature maps and the input image based on a preset graph attention network and a multi-channel attention mechanism to obtain the corresponding multi-view fused features. To enhance effective feature interaction and suppress information redundancy, this embodiment preferably first models the spatial correlation between different viewpoints through a graph attention network, and then uses a multi-channel attention mechanism based on tensor modality product to achieve adaptive weighted fusion of different features.
[0129] Specifically, the feature enhancement and fusion module includes a graph convolution processing unit, a channel stitching unit, an attention fusion unit, and a boundary enhancement unit connected in sequence. The graph convolution processing unit is used to construct a fully connected graph using the multi-view feature map and the input image as nodes, and to update the multi-view feature map and the input image according to the adjacency matrix of the fully connected graph and the preset graph attention network, to obtain the updated multi-view feature map and the updated input image. The construction of the fully connected graph can be implemented with reference to existing graph construction techniques, which will not be detailed here. The edge weights of the corresponding fully connected graph are obtained through adaptive learning by the graph attention network, and the update calculation process of the multi-view feature map and the input image (which can be understood as the feature map of the original acquisition viewpoint) is as follows:
[0130]
[0131] in, For nodes in a fully connected graph With nodes Edge weights between them; For activation functions; and This is a learnable weight parameter matrix; and They are nodes and nodes Feature map; Let be the adjacency matrix of a fully connected graph, and let the weight parameters between nodes be... ; For nodes The set of adjacent nodes One of the adjacent nodes; For graph attention networks; These are the updated vertical top-view feature map, diagonal oblique view feature map, side view feature map, and input image corresponding to the p-th input image, respectively.
[0132] The channel stitching unit is used to stitch the updated multi-view feature map and the updated input image together to obtain the corresponding joint feature tensor; wherein, the joint feature tensor is represented as:
[0133]
[0134] in, For channel splicing, It is a convolutional layer; Let be the joint feature tensor corresponding to the p-th input image.
[0135] The attention fusion unit can be understood as a processing module that adaptively allocates attention weight coefficients for features from different perspectives. It calculates the attention weight coefficients of each feature in the joint feature tensor based on a multi-channel attention mechanism using tensor modal product, and then fuses the updated multi-view feature map and the updated input image based on these attention weight coefficients to obtain the corresponding enhanced feature tensor. The calculation process of the enhanced feature tensor is as follows:
[0136] 1) A multi-channel attention mechanism based on tensor modal product adaptively captures joint feature tensors in each dimension. The weights of each feature and the attention weight coefficients of each feature are as follows:
[0137]
[0138] in, For activation function, This is the learnable parameter matrix based on tensor modal product used in the attention coefficient calculation. Represents the tensor and the learnable parameter matrix along the tensor's first... The operation of multiplying dimensions; These are the updated vertical top-view feature map, diagonal oblique view feature map, side view feature map, and attention weight coefficients of the input image, respectively.
[0139] 2) Multiply and sum the attention weight coefficients with their corresponding features to obtain the fused enhanced feature tensor:
[0140]
[0141] in, This is an element-wise multiplication operation; Let be the augmented feature tensor corresponding to the p-th input image.
[0142] The boundary enhancement unit can be understood as a dual-pooling channel boundary enhancement processing module used to further enhance feature boundary information. It performs average pooling and max pooling on the enhanced feature tensor, respectively, and then convolves the resulting first and second pooled features before fusing them with the enhanced feature tensor to obtain the multi-view fused feature. That is, for the enhanced feature tensor... Average pooling and max pooling are performed separately. Average pooling retains more background information to obtain the first pooling feature, while max pooling retains more texture information to obtain the second pooling feature. Then, convolution operations are performed on the first and second pooling features respectively to generate the corresponding two branch features. and Then, to prevent information loss, the features of the two branches are added together and then combined with the original enhanced feature tensor. Multiplying them together yields the final multi-view fused features. :
[0143]
[0144] In the formula, and These are max pooling and average pooling, respectively.
[0145] The defect target detection module is used to identify target defects based on the multi-view fusion features and a preset neural network to obtain the corresponding defect detection results. The preset neural network can use YOLOv8 as the backbone structure to complete the bounding box prediction and category classification of the target defect region, and obtain the defect detection results including the location coordinates of the defect target bounding box and the confidence score.
[0146] The first defect detection model with the above structure converts a single-view image into a multi-view image based on affine and perspective transformations, extracts features from different perspectives, and then fully explores the interaction relationships between features from different perspectives based on graph attention networks and multi-channel attention mechanisms. It also employs a multi-view feature fusion analysis mechanism that performs deep fusion and semantic enhancement of multi-view features. Compared to existing single-view recognition methods, this model can more comprehensively capture defect features under varying postures, complex structures, or partial occlusion, effectively improving the accuracy and comprehensiveness of defect identification and detection in complex inspection scenarios. It should be noted that the construction process of the first defect detection model in practical applications can be understood as follows: obtaining a model training set labeled with abnormal and normal line images of different defect targets; based on the model training set, using existing network model training methods, training and optimizing the network model with sequentially connected multi-view image conversion, multi-view feature extraction, feature enhancement and fusion, and defect target detection modules until the corresponding training termination conditions are met, thus obtaining the first defect detection model with stable model parameters.
[0147] In practical applications, a defect detection result can be obtained by performing multi-view fusion detection analysis on each frame image in the static visual image and dynamic inspection image dataset based on the first defect detection model in the cloud. To ensure that the cloud detection ultimately obtains a reliable target inspection result, this embodiment preferably performs a comprehensive analysis on the defect detection results corresponding to each frame image in the static visual image and dynamic inspection image dataset to determine the target inspection result. Specifically, the steps of the cloud performing defect detection based on the pre-built first defect detection model according to the static visual image and the dynamic inspection image dataset to obtain the target inspection result include:
[0148] Based on the first defect detection model, defect detection analysis is performed on each frame image in the static visual image and the dynamic inspection image dataset to obtain several defect detection results. The number of defect detection results corresponds to the total number of images involved in the static visual image and the dynamic inspection image dataset. The acquisition process of each defect detection result can refer to the data processing flow of each functional module in the aforementioned first defect detection model, which will not be repeated here.
[0149] When all the defect detection results indicate that there is no defective target, the target inspection result is set to normal; that is, when all defect detection results are normal, the inspection is considered normal.
[0150] When a defective target is found among all the defect detection results, the defect detection result with the highest confidence is taken as the target inspection result; that is, if at least one defect detection result indicates the presence of a defective target, the defect detection result with the highest confidence is taken as the final target inspection result.
[0151] This invention provides a method for acquiring static visual images of a target transmission line, generating corresponding static differential images based on these images, and performing local defect detection using the static differential images to obtain corresponding static detection results. When the static detection results meet a preset detection accuracy requirement, they are used as the target inspection results. When the static detection results do not meet the preset detection accuracy requirement, the static visual images are compressed and uploaded to the cloud, and a drone is triggered to perform an image re-acquisition task. The dynamic inspection image dataset acquired according to the image re-acquisition task is then compressed and uploaded to the cloud, enabling the cloud to use the static visual images and dynamic inspection images... The dataset is a technical solution for obtaining target inspection results by performing defect identification and analysis based on a pre-built first defect detection model for multi-view feature fusion detection of input images. On the basis of achieving efficient local defect detection by deploying static sensing equipment, it adopts a multi-view intelligent sensing architecture that combines dynamic and static elements and edge-cloud multi-source collaboration by using UAVs for emergency image supplementation and cloud-based multi-view fusion detection and analysis. This not only improves the spatial coverage and temporal continuity of line inspection, but also effectively enhances the efficiency and accuracy of line inspection, thereby improving the robustness and real-time response capability of intelligent inspection and meeting the intelligent inspection needs in complex power transmission environments.
[0152] It should be noted that although the steps in the flowchart above are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise explicitly stated in this document, there is no strict order requirement for the execution of these steps, and they can be executed in other orders.
[0153] In one embodiment, such as Figure 6As shown, a smart transmission line inspection system based on UAVs and cloud-edge collaboration is provided. The system includes:
[0154] The local detection module 1 is used to acquire a static visual image of the target transmission line, generate a corresponding static differential image based on the static visual image, and perform local defect detection based on the static differential image to obtain the corresponding static detection result.
[0155] The accuracy analysis module 2 is used to take the static detection result as the target inspection result when the static detection result meets the preset detection accuracy requirements.
[0156] The cloud-based detection module 3 is used to compress and upload the static visual image to the cloud when the static detection result does not meet the preset detection accuracy requirement, and to trigger the UAV to perform an image re-acquisition task, and to compress and upload the dynamic inspection image dataset collected according to the image re-acquisition task to the cloud, so that the cloud can perform defect identification analysis based on the static visual image and the dynamic inspection image dataset, based on the pre-built first defect detection model, to obtain the target inspection result; the first defect detection model is used to perform multi-view feature fusion detection on the input image.
[0157] Specific limitations regarding the intelligent transmission line inspection system based on UAVs and cloud-edge collaboration can be found in the above description of the intelligent transmission line inspection method based on UAVs and cloud-edge collaboration; the corresponding technical effects are equivalent and will not be repeated here. Each module in the aforementioned intelligent transmission line inspection system based on UAVs and cloud-edge collaboration can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0158] In summary, the intelligent inspection method and system for power transmission lines based on UAVs and cloud-edge collaboration provided by this invention, on the basis of local efficient defect detection by deploying static sensing equipment, adopts a dynamic and static combined with cloud-based multi-view fusion detection and analysis architecture that combines UAV emergency image acquisition with cloud-based multi-view fusion detection and analysis. This not only improves the spatial coverage and temporal continuity of line inspection, but also effectively enhances the efficiency and accuracy of line inspection, thereby improving the robustness and real-time response capability of intelligent inspection and meeting the intelligent inspection needs in complex power transmission environments.
[0159] The various embodiments in this specification are described in a progressive manner. For directly identical or similar parts of the embodiments, refer to each other. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0160] The above-described embodiments are merely preferred embodiments of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various improvements and substitutions without departing from the principles of the present invention, and these improvements and substitutions should also be considered within the scope of protection of the present invention. Therefore, the scope of protection of this invention should be determined by the scope of the claims.
Claims
1. A method for intelligent inspection of power transmission lines based on unmanned aerial vehicles (UAVs) and cloud-edge collaboration, characterized in that, The method includes: Acquire a static visual image of the target transmission line, generate a corresponding static differential image based on the static visual image, and perform local defect detection based on the static differential image to obtain the corresponding static detection result; When the static detection result meets the preset detection accuracy requirement, the static detection result is taken as the target inspection result; If the static detection result does not meet the preset detection accuracy requirement, the static visual image is compressed and uploaded to the cloud, and the UAV is triggered to perform an image re-acquisition task. The dynamic inspection image dataset collected according to the image re-acquisition task is also compressed and uploaded to the cloud, so that the cloud can perform defect identification analysis based on the static visual image and the dynamic inspection image dataset, based on the pre-built first defect detection model, to obtain the target inspection result; the first defect detection model is used to perform multi-view feature fusion detection on the input image; The step of performing local defect detection based on the static difference image to obtain the corresponding static detection result includes: The static difference image is input into a pre-constructed second defect detection model for defect identification and analysis to obtain the static detection result. The second defect detection model is used to adaptively enhance the texture edges of the static difference image based on its texture complexity, and then sequentially perform semantic feature extraction and defect target detection. The second defect detection model includes a dynamic routing module, a dual-channel edge enhancement module, a semantic feature extraction module, and a detection head module connected in sequence. The dual-channel edge enhancement module includes a first texture edge enhancement branch and a second texture edge enhancement branch in parallel. The dynamic routing module is used to calculate grayscale information entropy based on the proportion of pixels at different grayscale levels in the static difference image, and input the static difference image into the first texture edge enhancement branch or the second texture edge enhancement branch in the dual-channel edge enhancement module according to the relationship between the grayscale information entropy and the preset information entropy threshold. The dual-channel edge enhancement module is used to adaptively enhance the texture edges of the static difference image based on the texture complexity of the static difference image, and obtain the corresponding texture enhancement features. The semantic feature extraction module is used to extract multi-scale semantic features from the texture enhancement features to obtain the corresponding multi-scale semantic features; The detection head module is used to obtain the static detection result by sequentially performing global average pooling and fully connected processing on the multi-scale semantic features.
2. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 1, characterized in that, The step of generating a corresponding static difference image based on the static visual image includes: The static visual image is preprocessed to obtain a preprocessed static image; Based on the lighting conditions of the preprocessed static image, a corresponding static viewpoint reference image is obtained from a preset normal static visual image library. Pixel difference image and structure difference image are generated based on the preprocessed static image and the static viewpoint reference image; The pixel difference image and the structural difference image are weighted and fused to obtain the static difference image.
3. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 1, characterized in that, The first texture edge enhancement branch includes a dilated convolutional layer, a subpixel convolutional layer, and a channel attention layer connected in sequence; The second texture edge enhancement branch includes a depth convolutional layer, an attention layer, and a global pooling layer connected in sequence; the attention layer includes an improved window attention layer and an improved shifted window self-attention layer connected in sequence.
4. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 3, characterized in that, Both the improved window attention layer and the improved shift window self-attention layer modulate the attention weights based on the gradient of the feature map.
5. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 1, characterized in that, The first defect detection model includes a multi-view image conversion module, a multi-view feature extraction module, a feature enhancement and fusion module, and a defect target detection module connected in sequence. The multi-view image conversion module is used to perform multi-view image conversion on the input image to obtain a corresponding multi-view image set; the multi-view image set includes vertical top view image, diagonal oblique view image and side view image; The multi-view feature extraction module is used to extract features from the vertical top-view image based on a hollow spatial pyramid network, to extract features from the diagonal oblique-view image based on a deformable convolutional network, and to extract features from the side-view image based on vertical stripe pooling combined with horizontal convolution, to obtain corresponding multi-view feature maps; the multi-view feature maps include a vertical top-view feature map, a diagonal oblique-view feature map, and a side-view feature map. The feature enhancement and fusion module is used to perform feature enhancement and fusion on the multi-view feature map and the input image based on a preset graph attention network and a multi-channel attention mechanism to obtain the corresponding multi-view fused features. The defect target detection module is used to identify target defects based on the multi-view fusion features and a preset neural network to obtain the corresponding defect detection results.
6. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 5, characterized in that, The feature enhancement and fusion module includes a graph convolution processing unit, a channel splicing unit, an attention fusion unit, and a boundary enhancement unit connected in sequence. The graph convolution processing unit is used to construct a fully connected graph using the multi-view feature map and the input image as nodes, and to update the multi-view feature map and the input image according to the adjacency matrix of the fully connected graph and the preset graph attention network to obtain the updated multi-view feature map and the updated input image. The channel stitching unit is used to stitch the updated multi-view feature map and the updated input image together to obtain the corresponding joint feature tensor. The attention fusion unit is used to calculate the attention weight coefficients of each feature in the joint feature tensor based on the multi-channel attention mechanism of tensor modality product, and to fuse the updated multi-view feature map and the updated input image based on the attention weight coefficients of each feature to obtain the corresponding enhanced feature tensor. The boundary enhancement unit is used to perform average pooling and max pooling on the enhancement feature tensor, and then convolve the corresponding first pooling feature and second pooling feature, and then fuse them with the enhancement feature tensor to obtain the multi-view fused feature.
7. The intelligent inspection method for power transmission lines based on UAVs and cloud-edge collaboration as described in claim 1, characterized in that, The steps by which the cloud platform performs defect detection based on the static visual image and the dynamic inspection image dataset, using a pre-built first defect detection model, to obtain the target inspection result include: Based on the first defect detection model, defect detection analysis is performed on each frame image in the static visual image and the dynamic inspection image dataset to obtain several defect detection results. When all the defect detection results indicate that there is no defective target, the target inspection result is set to normal. When a defective target is found among all the defect detection results, the defect detection result with the highest confidence level is taken as the target inspection result.
8. A smart inspection system for power transmission lines based on unmanned aerial vehicles (UAVs) and cloud-edge collaboration, characterized in that, The system employing the intelligent transmission line inspection method based on UAVs and cloud-edge collaboration as described in claim 1 includes: The local detection module is used to acquire static visual images of the target transmission line, generate corresponding static differential images based on the static visual images, and perform local defect detection based on the static differential images to obtain corresponding static detection results. The accuracy analysis module is used to take the static detection result as the target inspection result when the static detection result meets the preset detection accuracy requirements. The cloud-based detection module is used to compress and upload the static visual image to the cloud when the static detection result does not meet the preset detection accuracy requirement, and to trigger the UAV to perform an image re-acquisition task. The module also compresses and uploads the dynamic inspection image dataset collected according to the image re-acquisition task to the cloud, so that the cloud can perform defect identification analysis based on the static visual image and the dynamic inspection image dataset, using a pre-built first defect detection model, to obtain the target inspection result. The first defect detection model is used to perform multi-view feature fusion detection on the input image.
Citation Information
Patent Citations
Dynamic and static cooperative power transmission line refined inspection method and system
CN115220479A
Power transmission line body defect detection method based on cloud edge cooperation
CN117975305A