Method and device for detecting stability of power transmission line tower footing and medium
By combining remote sensing technology and UAV inspections with a dual-channel feature fusion network model based on deep learning algorithms, the problem of low efficiency in traditional inspection methods has been solved. This enables rapid identification and real-time monitoring of the stability of transmission line tower foundations, reducing the impact of geological disasters on the power grid.
Patent Information
- Application Number
- CN202511274811.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2026-01-20
AI Technical Summary
Traditional methods for inspecting transmission line tower foundations are inefficient and costly, and it is difficult to quickly identify and assess the stability of tower foundations over a large area, especially when they are susceptible to geological disasters under conditions of heavy rainfall and flooding or prolonged immersion.
Remote sensing technology is used to acquire tower base image data, which is combined with UAV inspection data, and a dual-channel feature fusion network model is constructed using deep learning algorithms to achieve rapid identification and assessment of tower base stability.
It enables real-time monitoring and precise early warning over a wide area, allowing for timely detection of tower foundation stability issues, reducing the impact of geological disasters on transmission lines, minimizing power grid failures and outages, and reducing economic losses.
Smart Images

Figure CN121366348A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of power transmission line monitoring, in particular to a method and device for detecting the stability of a power transmission line tower foundation and a medium. BACKGROUND
[0002] With the intensification of climate change and the frequent occurrence of natural disasters, the stability of power transmission line tower foundations has become one of the important challenges faced by the power industry. In particular, under the conditions of heavy rainfall and long-term immersion, power transmission line tower foundations are prone to be affected by geological disasters, such as tower foundation tilting, foundation loosening, tower body cracking, etc., which may lead to the destruction and shutdown of power transmission lines, seriously affecting the stability and reliability of power supply.
[0003] Traditional methods of inspecting power transmission line tower foundations mainly rely on manual inspection, which has the problems of limited inspection range, low efficiency, high cost, etc. SUMMARY
[0004] The application provides a method, device and medium for detecting the stability of a power transmission line tower foundation, which has the advantage of obtaining image data and unmanned aerial vehicle inspection data of the power transmission line tower foundation through remote sensing technology, and analyzing and processing the data using a deep learning algorithm, which can realize rapid identification and evaluation of the stability of the tower foundation.
[0005] In one aspect, the application provides a method for detecting the stability of a power transmission line tower foundation based on remote sensing images and unmanned aerial vehicle cruising, comprising the following steps:
[0006] S1, data collection and dataset establishment:
[0007] Obtain remote sensing image data of the power transmission line tower foundation, obtain close-up image data of the power transmission line tower foundation by unmanned aerial vehicle cruising, and take the two groups of images as the original data of the dataset;
[0008] Establish a dataset according to the obtained data;
[0009] S2, construct a dataset containing labeled information and divide it into a training set, a validation set and a test set:
[0010] Label the data in the dataset to obtain a labeled file containing target frame information, and any image is matched with a corresponding labeled file;
[0011] Divide the dataset into a training set, a validation set and a test set;
[0012] S3, construct a dual-channel feature fusion network model and train it; the dual-channel feature fusion network model combines multi-scale feature enhancement, feature pyramid network, self-attention mechanism, residual connection and full convolution network modules, and realizes the detection of the stability of the power transmission tower foundation through a dual-channel structure;
[0013] S4, applying the trained dual-channel feature fusion network model to the tower foundation of the power transmission line for stability detection.
[0014] Further, in step S1, the acquired data is preprocessed and arranged, including removing outliers, filling missing values, and unifying data formats.
[0015] Further, in step S2, the remote sensing image and the close-range image taken by the unmanned aerial vehicle are labeled with tower foundation stability type labeling information; the tower foundation stability type labeling information includes three categories: stable, slight deviation, and serious deviation.
[0016] The labeling file includes class information and position information of the target frame, wherein the position information is composed of four data, namely the horizontal coordinate of the target frame center point, the vertical coordinate of the target frame center point, the target frame width, and the target frame height. The original power transmission line tower foundation stability image dataset and the corresponding labeling file constitute the power transmission line tower foundation stability image dataset containing labeling information.
[0017] Further, in step S3, the dual-channel feature fusion network model first sends the remote sensing image P1 of the power transmission line tower foundation collected by the satellite remote sensing in the data set to the backbone network ResNet-50 to obtain the feature map A, and sets the dimension of the output feature map as CxHxW, wherein C is the channel number, H and W are the height and width of the feature map. The feature map A of a specific third layer is obtained:
[0018]
[0019] wherein is the third layer convolution kernel parameter, and Conv represents the convolution operation.
[0020] Then, the obtained feature map A is input into the MSCA module to improve the multi-scale representation ability of the feature map A; the global channel information is extracted by using the global average pooling, and then the channel weight is generated by using the 1x1 convolution and the Sigmoid activation, and then multiplied with the original feature map A to generate the feature map B:
[0021]
[0022] wherein is the Sigmoid function, and then enters the batch normalization to standardize the feature distribution and stabilize the training process; the feature map C is obtained
[0023]
[0024] wherein, μ(B) and σ(B) are the mean and standard deviation of the feature map B, respectively, is a smooth factor to prevent zero, and γ and β are learnable parameters; the feature map C is put into the self-attention mechanism again, so that the feature map can capture long-range dependencies and enhance the sensitivity to global information, and a feature map D is obtained:
[0025]
[0026] where W q , W k and W v are the weight matrices of training, Q, K and V are query, key and value matrices, is the dimension of the key vector; then the feature map D enters the ReLU activation module, which helps the model express complex features by increasing nonlinearity, and a feature map E is obtained:
[0027]
[0028] The tower base image P2 taken by the unmanned aerial vehicle in the data set is input into the backbone network ResNet-50 to extract local structural features and obtain a third layer feature map F:
[0029]
[0030] where is the third layer convolution kernel parameter, and Conv represents convolution operation; the obtained feature map F is input into the FPN module, and FPN enhances the detail information of the feature map F at different scales, thereby obtaining a feature map G; in the FPN, the feature map F goes through several steps, including horizontal connection, upsampling and feature map fusion, and finally generates an output feature map G; for the input feature map F, it is converted into an intermediate feature map F1 through a 1×1 convolution layer:
[0031]
[0032] where Conv 1×1 represents 1×1 convolution operation; using layer-by-layer upsampling, the feature maps of lower layers are fused with the feature maps of higher layers; through the feature map F1 of the last layer, F up is obtained by upsampling, and then it is added to the feature map P of the current layer:
[0033]
[0034] UpSample represents the upsampling operation, and after step-by-step upsampling and fusion at each layer, a multi-scale feature map G is finally obtained, which includes low-resolution global information and high-resolution local information; then it enters the batch normalization to standardize the feature distribution; and a feature map H is obtained:
[0035]
[0036] where μ(G) and σ(G) are the mean and standard deviation of the feature map G, respectively, is a smoothing factor to prevent division by zero, and γ and β are learnable parameters; the feature map H is again passed through a self-attention mechanism to allow the feature map to capture long-range dependencies, resulting in feature map I, which is then passed through a ReLU activation module to increase nonlinearity and help the model express complex features, resulting in feature map J:
[0037]
[0038] The feature maps E and J generated by the two channels are connected in residual to fuse the features of channel one and channel two while preserving the original information:
[0039]
[0040] Subsequently, the fused features are further integrated and compressed through convolutional layers; generating feature map L:
[0041]
[0042] The resulting feature map L is then passed into the fully convolutional network FCN for further global analysis of the fused features. In the FCN, the input image L is passed through multiple convolutional layers to extract features, resulting in L conv , which is then passed through a ReLU activation to help the network introduce non-linear features and enable it to learn more complex mapping relationships, resulting in L relu . The resulting feature map L is then passed through a transpose convolutional layer for upsampling to restore the same spatial resolution as the input image L up . Finally, the Softmax function is applied to obtain the output M:
[0043]
[0044] W conv and b conv are the weights and biases of the convolutional layer, respectively, and W deconv and b deconv are the weights and biases of the transpose convolutional layer; finally, the generated feature map M is converted into class information and position information through the Softmax and MLP modules:
[0045]
[0046] where Wcls is the weight matrix of the Softmax layer, bcls is the bias vector of the Softmax layer, and MLPloc is the multi-layer perceptron for position prediction.
[0047] After the construction of the dual-channel feature fusion network is completed, the Hungarian algorithm is used to match the network output prediction target frame with the real target frame, and the prediction frame corresponding to the real target frame in the prediction target frame is selected; then, the classification loss and the position loss between the selected prediction frame and the real frame are calculated, and the sum of the two constitutes the total loss function;
[0048] The classification loss uses cross-entropy loss:
[0049]
[0050] wherein L cls represents the classification loss, represents the class label of the real frame, is the class probability of the prediction frame, and C is the total number of classes;
[0051] The position loss uses GIoU loss:
[0052]
[0053] wherein Ac is the area of the smallest closed region containing the real frame and the prediction frame; Au is the intersection area of the prediction frame and the real frame; IoU is the intersection-over-union of the real frame and the prediction frame; if the training loss is significantly lower than the validation loss, i.e., |d2 validation set loss-d1 training set loss|>0.05, it indicates that the model is too fitted on the training set and fails to effectively generalize, then at this time, regularization, data augmentation and other methods are adopted; if the training loss and the validation loss are both within an acceptable range and relatively stable, i.e., |d2 validation set loss-d1 training set loss|<0.05, then the dual-channel feature fusion network model training is completed.
[0054] Further, in step S4, the trained dual-channel feature fusion network model is configured to the terminal, and the terminal detects the tower group image of the power transmission line to be detected photographed by the unmanned aerial vehicle and the remote sensing image around the tower foundation photographed by the remote sensing satellite to obtain an output image with a target frame containing the result of whether the tower foundation is stable.
[0055] In another aspect, the present application provides a device for detecting the stability of a power transmission line tower foundation based on remote sensing images and unmanned aerial vehicle cruising, comprising a processor and a memory, wherein the memory stores a computer program, and the computer program is called and executed by the processor to implement the method for detecting the stability of a power transmission line tower foundation based on remote sensing images and unmanned aerial vehicle cruising as described above.
[0056] In another aspect, the present application provides a computer readable medium storing a computer program, and the computer program is called and executed by a computer to implement the method for detecting the stability of a power transmission line tower foundation based on remote sensing images and unmanned aerial vehicle cruising as described above.
[0057] In summary, the application has the following advantages:
[0058] 1. The application can be used to detect tower foundation stability problems of power transmission lines in a large-scale range in a timely manner by using remote sensing images and data obtained by unmanned aerial vehicle cruising, and using a dual-channel feature fusion network model to identify and detect image data.
[0059] 2. The application can perform real-time monitoring using remote sensing technology, unmanned aerial vehicle inspection, and sensor monitoring to achieve real-time monitoring of geological disasters and tower foundation stability of power transmission lines, timely detection of problems, and timely emergency measures. At the same time, the detection results can be accurately warned, and the possibility of geological disasters can be accurately predicted by combining prediction models and data analysis techniques, and timely warning information can be issued to prepare in advance. Moreover, it can reduce losses by early warning and timely response to reduce the impact of geological disasters on power transmission lines, reduce power grid failures and outages, and thus reduce economic losses and social impact. BRIEF DESCRIPTION OF DRAWINGS
[0060] Figure 1 is a flowchart of a method for detecting tower foundation stability of a power transmission line based on remote sensing images and unmanned aerial vehicle cruising in an embodiment of the application;
[0061] Figure 2 is a network structure diagram of a dual-channel feature fusion network model in an embodiment of the application. DETAILED DESCRIPTION
[0062] The specific embodiments of the application will be described in detail below with reference to the accompanying drawings.
[0063] A specific embodiment of the application provides a method for detecting tower foundation stability of a power transmission line based on remote sensing images and unmanned aerial vehicle cruising, as shown in Figure 1 The method comprises the following steps:
[0064] S1, data collection and data set establishment:
[0065] Obtain remote sensing image data of the tower foundation of the power transmission line, obtain close-range image data of the tower foundation of the power transmission line by unmanned aerial vehicle cruising, and use the two groups of images as the original data of the data set; preprocess and organize the obtained data, including removing outliers, filling missing values, and unifying data formats to ensure data quality and usability.
[0066] According to the obtained data, a data set is established, which provides a reliable data basis for training of a new model.
[0067] S2, construct a data set containing labeled information and divide the training set, validation set, and test set:
[0068] The data in the data set is labeled to obtain a label file containing target frame information, and any image is matched with a corresponding label file;
[0069] The object requiring the label information is a remote sensing image and a close-range image shot by a drone. The remote sensing image and the close-range image shot by the drone are labeled with tower foundation stability type label information. The tower foundation stability type label information includes three categories of stable, slight deviation and serious deviation;
[0070] The label file includes category information and position information of the target frame. The position information is composed of four data, i.e., a horizontal coordinate of a center point of the target frame, a vertical coordinate of the center point of the target frame, a width of the target frame and a height of the target frame. The original power transmission line tower foundation stability image data set and the corresponding label file constitute a power transmission line tower foundation stability image data set containing label information.
[0071] The data set is divided into a training set, a validation set and a test set;
[0072] S3, a dual-channel feature fusion network model is constructed and trained. The dual-channel feature fusion network model combines a multi-scale feature enhancement, a feature pyramid network, a self-attention mechanism, a residual connection and a full convolution network module, and realizes detection of power tower foundation stability through a dual-channel structure;
[0073] The dual-channel feature fusion network model can be applied to prediction and analysis of power transmission line tower foundation stability detection, as shown in Figure 2 The network structure is a multi-stage processing flow, and the structure flow is as follows:
[0074] The dual-channel feature fusion network model first sends a remote sensing image P1 of a power transmission tower foundation collected by satellite remote sensing in the data set to a backbone network ResNet-50. The residual connection enables ResNet to maintain effective gradient descent in a very deep network, is suitable for ladder feature extraction tasks from simple to complex, and is suitable for multi-pixel image extraction. A feature map A is obtained. The dimension of the output feature map is set to CxHxW, where C is the number of channels, and H and W are the height and width of the feature map. A feature map A of a specific third layer is obtained:
[0075]
[0076] wherein represents the output feature map extracted at the second layer of the ResNet network, is a third layer convolution kernel parameter, and Conv represents a convolution operation;
[0077] The obtained feature map A is then input into the MSCA module to improve the multi-scale representation capability of the feature map A; global channel information is extracted using global average pooling (GAP), channel weights are generated through 1x1 convolution and Sigmoid activation, and then multiplied with the original feature map A to generate feature map B:
[0078]
[0079] wherein is a Sigmoid function, and then enters batch normalization to regulate feature distribution and stabilize the training process; the feature map C is obtained
[0080]
[0081] wherein μ(B) and σ(B) are the mean and standard deviation of the feature map B respectively, is a smoothing factor to prevent division by zero, and γ and β are learnable parameters; the feature map C is again subjected to self-attention mechanism to enable the feature map to capture long-range dependencies and enhance sensitivity to global information, and the feature map D is obtained:
[0082]
[0083] wherein W q , W k and W v are trained weight matrices, Q, K and V are query, key and value matrices respectively, is the dimension of the key vector; the feature map D then enters the ReLU activation module to increase nonlinearity and help the model express complex features, and the feature map E is obtained:
[0084]
[0085] The tower base image P2 taken by a drone in the data set is fed into the backbone network ResNet-50 to extract local structural features and obtain the third layer feature map F:
[0086]
[0087] wherein is the third layer convolution kernel parameter, and Conv represents convolution operation; the obtained feature map F is input into the FPN module, and FPN enhances the detail information of the feature map F at different scales to obtain the feature map G; in the FPN, the feature map F undergoes several steps including horizontal connection, up-sampling and feature map fusion to finally generate the output feature map G; for the input feature map F, it is converted into an intermediate feature map F1 through a 1x1 convolution layer:
[0088]
[0089] where Conv 1×1 denotes a 1x1 convolution operation; the feature maps of lower layers are fused with the feature maps of high layers using layer-by-layer upsampling; F up is obtained by upsampling through the feature maps F1 of the previous layer, and then added to the feature maps P of the current layer:
[0090]
[0091] UpSample denotes an upsampling operation. After step-by-step upsampling and fusion at each layer, a multi-scale feature map G is finally obtained, which includes low-resolution global information and high-resolution local information; then it enters the batch normalization to standardize the feature distribution; and a feature map H is obtained:
[0092]
[0093] where μ(G) and σ(G) are the mean and standard deviation of the feature map G, is a smoothing factor to prevent division by zero, and γ and β are learnable parameters; the feature map H is again subjected to the self-attention mechanism, allowing the feature map to capture long-range dependencies, and a feature map I is obtained; then it enters the ReLU activation module to increase nonlinearity, helping the model to express complex features, and a feature map J is obtained:
[0094]
[0095] The feature maps E and J generated by the two channels are connected in residual, which is used to fuse the features of channel one and channel two while preserving the original information:
[0096]
[0097] Subsequently, it enters the convolution layer to further integrate and compress the fused features; and a feature map L is generated:
[0098]
[0099] The obtained feature map L enters the fully convolutional network FCN for further global analysis of the fused features. In the FCN, the input image L is subjected to feature extraction through multiple convolution layers to obtain L conv , and ReLU activation is used to help the network introduce nonlinear features, enabling it to learn more complex mapping relationships to obtain L relu . Then, the transposed convolution layer is used for upsampling to restore the spatial resolution to the same as the input image L up , and finally the Softmax function is applied to obtain the output M:
[0100]
[0101] W conv and b conv These are the weights and biases of the convolutional layer, W. deconv and b deconv These are the weights and biases of the transposed convolutional layer; finally, the generated feature map M is converted into category information and location information by the Softmax and MLP modules respectively.
[0102]
[0103] Among them, W cls This is the weight matrix of the Softmax layer, b cls It is the bias vector of the Softmax layer, MLP loc It is a multilayer perceptron for location prediction;
[0104] After completing the construction of the dual-channel feature fusion network, the Hungarian algorithm is used to match the network output predicted target boxes with the real target boxes, and select the predicted target boxes that correspond to the real target boxes. Then, the classification loss and position loss between the selected predicted boxes and the real boxes are calculated, and the two are added together to form the total loss function.
[0105] The classification loss uses cross-entropy loss:
[0106]
[0107] Among them, L cls Represents classification loss, The category label representing the real bounding box. is the probability of the predicted bounding box category, and C is the total number of categories;
[0108] Location loss uses GIoU loss:
[0109]
[0110] Where Ac is the area of the smallest closure region containing the ground truth bounding box and the predicted bounding box; Au is the intersection area of the predicted bounding box and the ground truth bounding box; IoU is the intersection-union ratio of the ground truth bounding box and the predicted bounding box; if the training loss is significantly lower than the validation loss, i.e., |d2 validation loss - d1 training loss| > 0.05, it indicates that the model is overfitting on the training set and has failed to generalize effectively. In this case, regularization, data augmentation and other methods are adopted. If both the training loss and the validation loss are within an acceptable range and relatively stable, i.e., |d2 validation loss - d1 training loss| < 0.05, then the training of the dual-channel feature fusion network model is complete.
[0111] S4. The trained dual-channel feature fusion network model is used to perform stability detection on the transmission tower base.
[0112] The trained dual-channel feature fusion network model is configured to a terminal, and an image of a tower group to be detected photographed by a UAV and a remote sensing image around a tower foundation photographed by a remote sensing satellite are detected in the terminal to obtain an output image with a target frame containing a result of whether the tower foundation is stable.
[0113] Another specific embodiment of the present application provides a device for detecting stability of a tower foundation of a power transmission line based on remote sensing images and UAV cruising, comprising a processor and a memory, wherein the memory stores a computer program, and the computer program is called and executed by the processor to implement the method for detecting stability of the tower foundation of the power transmission line based on the remote sensing images and the UAV cruising.
[0114] Another specific embodiment of the present application provides a computer readable medium storing a computer program, and the computer program is called and executed by a computer to implement the method for detecting stability of a tower foundation of a power transmission line based on remote sensing images and UAV cruising.
[0115] The above only describes the preferred embodiments of the present application, and it should be noted that, for those skilled in the art, without departing from the creative concept of the present application, several modifications and improvements can be made, which are all within the protection scope of the present application.
Claims
1. A method for detecting the stability of transmission line tower foundations based on remote sensing images and UAV patrols, characterized in that, Includes the following steps: S1, Data Collection and Dataset Establishment: Remote sensing image data of the transmission line tower base was acquired, and close-up image data of the transmission line tower base was acquired through drone cruise photography. The two sets of images were used as the raw data of the dataset. Create a dataset based on the obtained data; S2, Construct a dataset with labeled information and divide it into training, validation, and test sets: The data in the dataset is labeled to obtain a label file containing target bounding box information, and any image is matched with a corresponding label file; The dataset is divided into a training set, a validation set, and a test set; S3, Construct and train a dual-channel feature fusion network model; The dual-channel feature fusion network model combines multi-scale feature enhancement, feature pyramid network, self-attention mechanism, residual connection and fully convolutional network module to achieve stability detection of transmission tower base through dual-channel structure; S4. The trained dual-channel feature fusion network model is used to perform stability detection on the transmission tower base.
2. The method for detecting the stability of transmission line tower foundations based on remote sensing images and UAV patrols according to claim 1, characterized in that, In step S1, the acquired data is preprocessed and organized, including removing outliers, filling in missing values, and standardizing the data format.
3. The method for detecting the stability of transmission line tower foundations based on remote sensing images and UAV patrols according to claim 1, characterized in that, In step S2, the stability type of the tower base is labeled on the remote sensing image and the close-up image taken by the UAV; the stability type of the tower base includes three categories: stable, slightly offset, and severely offset. The annotation file includes category information and location information of the target bounding box. The location information consists of four data points: the x-coordinate of the target bounding box center point, the y-coordinate of the target bounding box center point, the width of the target bounding box, and the height of the target bounding box. The original transmission line tower foundation stability image dataset and the corresponding annotation file constitute the transmission line tower foundation stability image dataset containing annotation information.
4. The method for detecting the stability of transmission line tower foundations based on remote sensing images and UAV patrols according to claim 1, characterized in that, In step S3, the dual-channel feature fusion network model first feeds the remote sensing image P1 of the power transmission tower base collected by satellite remote sensing in the dataset into the backbone network ResNet-50 to obtain feature map A. The dimensions of the output feature map are set to C×H×W, where C is the number of channels, and H and W are the height and width of the feature map, thus obtaining the feature map A of the specific third layer: ; in These are the parameters of the third convolutional kernel; Conv represents the convolution operation. The resulting feature map A is then input into the MSCA module to enhance its multi-scale representation capability. Global average pooling is used to extract global channel information, and channel weights are generated through 1×1 convolution and sigmoid activation. These weights are then multiplied by the original feature map A to generate feature map B. ; in The Sigmoid function is used, and then batch normalization is performed to standardize the feature distribution and stabilize the training process. Feature map C is obtained ; Where μ(B) and σ(B) are the mean and standard deviation of feature map B, respectively. γ and β are smoothing factors to prevent division by zero; these are learnable parameters. Feature map C is then passed through a self-attention mechanism again to capture long-range dependencies and enhance its sensitivity to global information, resulting in feature map D. ; Among them W q W k and W v This is the training weight matrix, where Q, K, and V are the query, key, and value matrices, respectively. This is the dimension of the key vector; subsequently, the feature map D enters the ReLU activation module, where non-linearity is added to help the model express complex features, resulting in the feature map E: ; The tower base image P2 taken by the drone in the dataset is fed into the backbone network ResNet-50 to extract local structural features and obtain the third layer feature map F: ; in These are the parameters of the third convolutional kernel, where Conv represents the convolution operation. The resulting feature map F is fed into the FPN module. The FPN enhances the detailed information of the feature map F at different scales, thus obtaining the feature map G. In the FPN, the feature map F undergoes several steps, including lateral concatenation, upsampling, and feature map fusion, ultimately generating the output feature map G. For the input feature map F, a 1×1 convolutional layer is used to transform it into an intermediate feature map F1. ; Where Conv 1×1 This represents a 1×1 convolution operation; layer-by-layer upsampling is used to fuse the feature maps of lower layers with those of higher layers; F is obtained by upsampling the feature map F1 from the previous layer. up Then add it to the feature map P of the current layer: ; UpSample represents the upsampling operation. After progressive upsampling and fusion at each layer, a multi-scale feature map G is finally obtained. These feature maps include low-resolution global information and high-resolution local information. Then, batch normalization is performed to standardize the feature distribution, resulting in the feature map H. ; Where μ(G) and σ(G) are the mean and standard deviation of the feature map G, respectively. γ and β are smoothing factors to prevent division by zero; they are learnable parameters. The feature map H is then passed through a self-attention mechanism to capture long-range dependencies, resulting in feature map I. This I then enters the ReLU activation module, where non-linearity is added to help the model express complex features, resulting in feature map J. ; The feature maps E and J generated from the two channels are residually concatenated to fuse the features of channel one and channel two while preserving the original information. ; The features are then further integrated and compressed in a convolutional layer, generating feature map L. ; The resulting feature map L is fed into a fully convolutional network (FCN) for further global analysis of the fused features. Within the FCN, the input image L undergoes feature extraction through multiple convolutional layers to obtain L0. conv Furthermore, ReLU activation is used to introduce non-linear features into the network, enabling it to learn more complex mapping relationships to obtain L. relu Then, upsampling is performed through a transposed convolutional layer to restore the spatial resolution L to the same as the input image. up Finally, the Softmax function is applied to obtain the output M: ; W conv and b conv These are the weights and biases of the convolutional layer, W. deconv and b deconv These are the weights and biases of the transposed convolutional layer; finally, the generated feature map M is converted into category information and location information by the Softmax and MLP modules respectively. ; Where Wcls is the weight matrix of the Softmax layer, bcls is the bias vector of the Softmax layer, and MLPloc is the multilayer perceptron for location prediction. After completing the construction of the dual-channel feature fusion network, the Hungarian algorithm is used to match the network output predicted target boxes with the real target boxes, and select the predicted target boxes that correspond to the real target boxes. Then, the classification loss and position loss between the selected predicted boxes and the real boxes are calculated, and the two are added together to form the total loss function. The classification loss uses cross-entropy loss: ; Among them, L cls Represents classification loss, The category label representing the real bounding box. is the probability of the predicted bounding box category, and C is the total number of categories; Location loss uses GIoU loss: ; Where Ac is the area of the smallest closure region containing the ground truth bounding box and the predicted bounding box; Au is the intersection area of the predicted bounding box and the ground truth bounding box; IoU is the intersection-union ratio of the ground truth bounding box and the predicted bounding box; if the training loss is significantly lower than the validation loss, i.e., |d2 validation loss - d1 training loss| > 0.05, it indicates that the model is overfitting on the training set and has failed to generalize effectively. In this case, regularization, data augmentation and other methods are adopted. If both the training loss and the validation loss are within an acceptable range and relatively stable, i.e., |d2 validation loss - d1 training loss| < 0.05, then the training of the dual-channel feature fusion network model is complete.
5. The method for detecting the stability of transmission line tower foundations based on remote sensing images and UAV patrols according to claim 1, characterized in that, In step S4, the trained dual-channel feature fusion network model is configured on the terminal. The terminal detects the images of the transmission line towers to be detected taken by the UAV and the remote sensing images of the area around the tower base taken by the remote sensing satellite, and obtains an output image with a target box containing the result of whether the tower base is stable.
6. A device for detecting the stability of transmission line tower foundations based on remote sensing images and UAV patrols, characterized in that, It includes a processor and a memory, the memory storing a computer program, which, when executed by the processor, implements the method for detecting the stability of transmission line tower foundations based on remote sensing images and UAV cruise as described in claims 1-5.
7. A computer-readable medium, characterized in that, The computer-readable medium stores a computer program, which, when executed by a computer, implements the method for detecting the stability of transmission line tower foundations based on remote sensing images and UAV patrols as described in claims 1-5.