High-resolution satellite remote sensing image power transmission tower extraction method
By building a small-objective extraction model, using the LSKA-AFF module to enhance feature extraction and P-PANet network for multi-scale feature fusion, the problem of transmission tower detection in complex backgrounds in high-resolution satellite remote sensing images is solved, and higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202510340584.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-21
AI Technical Summary
In high-resolution satellite remote sensing images, traditional deep learning object detection methods are difficult to effectively identify transmission towers in complex backgrounds. This is mainly because the feature extraction network has a small proportion of features in small object detection, which makes it difficult for the network to identify, and the semantic information of high-level features may be lost or degraded under the challenges of multi-scale detection.
A high-resolution satellite remote sensing image transmission tower extraction method is proposed to build a small-objective extraction model, including feature extraction network, feature fusion network and object detection head. The feature extraction network enhances the target area characteristics through the LSKA-AFF module and expands the receptive field; the feature fusion network integrates multi-scale feature fusion through the progressive feature pyramid network and the path aggregation network to enhance the detection ability of targets at different scales.
Effectively capture global and local features, solve the problem of uneven and generally small target sizes in complex backgrounds, improve the model's detection ability in complex backgrounds, and significantly improve the detection accuracy of the target to be extracted in high-resolution satellite remote sensing images.
Smart Images

Figure CN119992366A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of image recognition, and in particular relates to a method for extracting transmission towers from high-resolution satellite remote sensing images. Background Art
[0002] Transmission line inspection is an important means to ensure the stable operation of the power system. High-resolution satellite remote sensing technology makes non-contact, large-scale environmental and infrastructure monitoring possible, providing a new solution for transmission line inspection. Transmission tower detection is the basis of transmission line inspection.
[0003] Traditional optical satellite remote sensing image target detection methods usually rely on machine learning and image processing technology, involving key steps such as region selection, feature extraction and classifier design. However, the detection model based on traditional image processing methods has poor generalization ability in different scenarios. Deep learning methods have shown certain effectiveness in detecting power towers in satellite remote sensing images, but it is still difficult to directly apply to the detection task of transmission towers in complex backgrounds in high-resolution satellite images. The reasons are:
[0004] In the target detection method of deep learning, the feature extraction network is used to extract high-dimensional features. However, in the detection of small targets in satellite remote sensing images, the extracted features account for a small proportion of the output feature map, making it difficult for the network to recognize;
[0005] Transmission tower detection in satellite remote sensing images faces the challenge of multi-scale detection. In the baseline model, high-level features need to be propagated and interacted at multiple intermediate scales before they can be fused with the underlying features. During this propagation and interaction process, the semantic information of high-level features may be lost or degraded. Summary of the invention
[0006] In view of this, the present invention aims to propose a method for extracting transmission towers from high-resolution satellite remote sensing images, in order to solve at least one of the above-mentioned technical problems.
[0007] To achieve the above object, the technical solution of the present invention is achieved as follows:
[0008] A method for extracting transmission towers from high-resolution satellite remote sensing images, comprising the following steps:
[0009] Acquire satellite remote sensing images with target location information to be extracted;
[0010] Constructing a small target extraction model, the small target extraction model consists of a feature extraction network, a feature fusion network, and a target detection head; wherein the feature extraction network captures feature maps from satellite remote sensing images and weights the feature maps through an attention mechanism, and the feature fusion network performs feature fusion on the weighted feature maps;
[0011] Train the small object extraction model;
[0012] The trained small target extraction model is used to extract the target information to be extracted from the real-time satellite remote sensing image, and the precise monitoring and management of the power system is completed based on the target information to be extracted.
[0013] Furthermore, the process of obtaining the satellite remote sensing image with the target location information to be extracted is as follows:
[0014] Acquire multiple high-resolution satellite remote sensing images of the target to be extracted, and combine the multiple satellite remote sensing images into a data set;
[0015] The locations of the targets to be extracted in all images of the data set are annotated, and the annotated data set is divided into a validation set and a training set in a ratio of 1:9.
[0016] Furthermore, the feature extraction network is composed of multiple convolutional layers and LSKA-AFF modules. The satellite remote sensing images processed by the multiple convolutional layers are used as the original feature map, and the following operations are performed:
[0017] The original feature map is sequentially passed through the batch normalization layer, 1×1 convolution, Gaussian error linear unit, LSKA module, and 1×1 convolution to obtain the first feature map;
[0018] The original feature map and the first feature map are passed into the AFF module for processing to obtain a second feature map;
[0019] The second feature map is sequentially passed through a batch normalization layer and a convolutional feed-forward layer to obtain a third feature map;
[0020] The second feature map and the third feature map are passed into the AFF module for processing to obtain the output feature map.
[0021] Furthermore, the LSKA module receives the feature map output by the Gaussian error linear unit and performs the following operations:
[0022] The feature map is convolved with the depth convolution after two layers of splitting to get the output
[0023] Will The convolution operation is performed with the depth-wise separable convolution after the two layers are split, and the output A is obtained after 1×1 convolution C ;
[0024] Output feature maps F and A C The element-wise product of .
[0025] Furthermore, the AFF module receives two feature maps and performs the following operations:
[0026] The weight array is obtained by summing the elements of the two feature maps;
[0027] Performing an element-wise product of a weight array and a feature map, and performing an element-wise product of a reverse array and another feature map, wherein the reverse array = 1-weight array;
[0028] Add the results of the two element-wise multiplications to obtain the weighted feature map.
[0029] Furthermore, the feature fusion network is composed of a progressive feature pyramid network and a path aggregation network. The working process of the feature fusion network is as follows:
[0030] Get features from multiple convolutional layers of the feature extraction network;
[0031] Input low-level features into the front progressive feature pyramid network;
[0032] The output of the front progressive feature pyramid network and the high-level features are input into the rear progressive feature pyramid network to obtain multi-scale features;
[0033] The multi-scale features are input into the path aggregation network to obtain the detection results of the target to be extracted.
[0034] Furthermore, the P-PANet network includes an ASFF module, which receives the weighted feature map and performs the following operations:
[0035] Extract feature vectors at the same position in feature maps with different weights;
[0036] Re-adjust the weight corresponding to each eigenvector until the weights of all eigenvectors meet the constraint condition, where the sum of all weights is 1;
[0037] Perform weighted fusion on the feature vectors corresponding to each weight.
[0038] Furthermore, during the training of the small target extraction model:
[0039] The batch size of the GPU is set to 16; the decay strategy selects cosine annealing; the optimizer selects the stochastic gradient descent algorithm; the initial learning rate is set to 0.01; the dynamic parameter is set to 0.937; the number of training rounds is set to train 350 traversals; the performance indicators of the small target extraction model are monitored, and when the performance indicators do not improve significantly within 100 consecutive traversals, the training is terminated and completed.
[0040] Before further extracting the target information to be extracted from the real-time satellite remote sensing image, the performance indicators of the small target extraction model are verified. If the verification result is up to standard, the small target extraction model is put into use. Otherwise, the small target extraction model is adjusted until the performance indicators of the small target extraction model meet the standards.
[0041] Furthermore, the performance indicators include:
[0042] Precision: the proportion of samples predicted as positive by the small object extraction model that are correctly predicted;
[0043] Recall rate is the proportion of samples that are actually positive and are correctly predicted by the small target extraction model;
[0044] F1 score, the harmonic mean of precision and recall;
[0045] mAP0.5, used to measure the accuracy of the model in the detection task.
[0046] Compared with the prior art, the method for extracting transmission towers from high-resolution satellite remote sensing images described in the present invention has the following beneficial effects:
[0047] The feature extraction network dynamically enhances the features of the target area, expands the receptive field, and effectively captures global and local features, solving the problem of uneven and generally small target sizes in complex backgrounds. Through multi-scale feature fusion, the detection capability of targets of different scales is enhanced, thereby improving the performance of the model in complex backgrounds. The feature fusion network further improves the model's target detection capability in complex backgrounds through gradual feature fusion and adaptive spatial weighting. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings:
[0049] Figure 1 A schematic diagram of a process of extracting a target to be extracted from a high-resolution satellite remote sensing image according to an embodiment of the present invention;
[0050] Figure 2 Schematic diagram of the LSKAFF-YOLO structure according to an embodiment of the present invention;
[0051] Figure 3 This is a schematic diagram of the LSKA-AFF structure according to an embodiment of the present invention;
[0052] Figure 4 This is a schematic diagram of the LSKA structure according to an embodiment of the present invention;
[0053] Figure 5 AFF structure schematic diagram according to an embodiment of the present invention;
[0054] Figure 6 A schematic diagram of the P-PANet structure according to an embodiment of the present invention;
[0055] Figure 7 This is a schematic diagram of the ASFF structure described in an embodiment of the present invention. DETAILED DESCRIPTION
[0056] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0057] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and the like are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, features defined as "first", "second", and the like may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.
[0058] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood by specific circumstances.
[0059] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0060] like Figure 1 As shown, a method for extracting transmission towers from high-resolution satellite remote sensing images comprises the following steps:
[0061] S1. Obtain a satellite remote sensing image with target location information to be extracted.
[0062] The target to be extracted is a transmission tower. The process of obtaining a satellite remote sensing image with the location information of the target to be extracted in step S1 is as follows:
[0063] Acquire multiple high-resolution satellite remote sensing images of the target to be extracted, and combine the multiple satellite remote sensing images into a data set;
[0064] The locations of the targets to be extracted in all images of the data set are annotated, and the annotated data set is divided into a validation set and a training set in a ratio of 1:9.
[0065] In some embodiments, the specific execution process of step S1 is as follows:
[0066] A dataset is composed of at least 3,000 high-resolution satellite images of the target to be extracted. The locations of power towers are marked with the Labelme tool in the images in the dataset. All data are divided into a training set and a validation set in a ratio of 9:1. In this way, a labeled dataset is created to form a training set and a test set for training the model.
[0067] S2. Construct a small target extraction model, wherein the small target extraction model consists of a feature extraction network, a feature fusion network, and a target detection head;
[0068] The feature extraction network captures feature maps from satellite remote sensing images and weights the feature maps through an attention mechanism, and the feature fusion network performs feature fusion on the weighted feature maps.
[0069] like Figure 2 As shown, in some embodiments, the small target extraction model is specifically a LSKAFF-YOLO network, and the LSKAFF-YOLO network is composed of a feature extraction network, a feature fusion network, and a target detection head. The specific parameters of the network are shown in the following table:
[0070]
[0071]
[0072] The feature extraction network in step S2 is composed of multiple convolutional layers and LSKA-AFF modules. The satellite remote sensing image processed by multiple convolutional layers is used as the original feature map, and the following operations are performed:
[0073] The original feature map is sequentially passed through the batch normalization layer, 1×1 convolution, Gaussian error linear unit, LSKA module, and 1×1 convolution to obtain the first feature map;
[0074] The original feature map and the first feature map are passed into the AFF module for processing to obtain a second feature map;
[0075] The second feature map is sequentially passed through a batch normalization layer and a convolutional feed-forward layer to obtain a third feature map;
[0076] The second feature map and the third feature map are passed into the AFF module for processing to obtain the output feature map.
[0077] The LSKA module receives the feature map output by the Gaussian error linear unit and performs the following operations:
[0078] The feature map is convolved with the depth convolution after two layers of splitting to get the output
[0079] Will The convolution operation is performed with the depth-wise separable convolution after the two layers are split, and the output A is obtained after 1×1 convolution C ;
[0080] Output feature maps F and A C The element-wise product of .
[0081] The AFF module receives two feature maps and performs the following operations:
[0082] The weight array is obtained by summing the elements of the two feature maps;
[0083] Performing an element-wise product of a weight array and a feature map, and performing an element-wise product of a reverse array and another feature map, wherein the reverse array = 1-weight array;
[0084] Add the results of the two element-wise multiplications to obtain the weighted feature map.
[0085] In some embodiments, Figure 3 As shown, an LSKA-AFF structure is constructed in the feature extraction network. The structure has 9 layers for enhancing the features of the target to be extracted. The module is mainly composed of a BN (batch normalization) layer, an LSKA module, an AFF module, a GELU (Gaussian error linear unit) and a CFFN (convolutional feedforward) layer.
[0086] The structure of the LSKA module is as follows Figure 4 As shown, the output of the LSKA module is shown below. It combines large convolution kernels and attention mechanisms to significantly expand the receptive field of the model, capture long-distance spatial dependencies, enhance the ability to extract global and local features of the target area to be extracted, and reduce the misidentification of background noise and complex objects.
[0087] Given an input feature map F∈R C×H×W , where C is the number of input channels, H and W represent the height and width of the feature map respectively. The output of LSKA can be obtained by formulas (1) to (4). The input and output calculation formulas are as follows:
[0088]
[0089] A C =W 1×1 *Z C , (3)
[0090]
[0091] Among them, the symbols * and Represent the convolution operation and element-wise product respectively;
[0092] The d in formula (1) represents the void ratio. By convolving the feature map F with two layers of convolution kernel W with a size of (2d-1), the output of DW-Conv can be obtained: This convolution is used to capture local spatial information and compensate for the network effect in the subsequent two-layer decomposed DW-D-Conv (Formula (2)).
[0093] It should be noted that each channel C of the feature map F will be convolved with the corresponding channel of the convolution kernel W, and then the size of the convolution kernel in DW-D-Conv is in Represents a floor operation, which is responsible for capturing The global spatial information of the convolution kernel W is represented by k, where k also represents the maximum receptive field of the convolution kernel W, which increases the model's ability to distinguish different features and effectively handles complex backgrounds and targets.
[0094] The output of the DW-D-Conv layer is finely convolved through a 1×1 convolution kernel to obtain the attention map A. C , the output of LSKA is the attention map A C And the input feature map F C The element-wise product of .
[0095] The structure of the AFF module is as follows Figure 5 As shown in Figure 2, the AFF module is used as a feature fusion module at the connection, focusing on improving the context information aggregation capability and feature fusion quality;
[0096] The AFF module can optimize the distribution of feature weights by adjusting the attention weights in the channel dimension, thereby extracting target features more accurately and suppressing irrelevant background information. Its input and output formulas are as follows:
[0097]
[0098] in, The feature fusion method representing the sum of elements, represents element-wise product;
[0099] It should be noted that the fusion weight It consists of real numbers between 0 and 1. Likewise, this enables the network to do a soft selection or weighted average between X and Y, which are two different feature maps.
[0100] The feature fusion network described in step S2 is composed of a progressive feature pyramid network and a path aggregation network. The working process of the feature fusion network is as follows:
[0101] Get features from multiple convolutional layers of the feature extraction network;
[0102] Input low-level features into the front progressive feature pyramid network;
[0103] The output of the front progressive feature pyramid network and the high-level features are input into the rear progressive feature pyramid network to obtain multi-scale features;
[0104] The multi-scale features are input into the path aggregation network to obtain the detection results of the target to be extracted.
[0105] The feature fusion network includes an ASFF module, which receives the weighted feature map and performs the following operations:
[0106] Extract feature vectors at the same position in feature maps with different weights;
[0107] Re-adjust the weight corresponding to each eigenvector until the weights of all eigenvectors meet the constraint condition, where the sum of all weights is 1;
[0108] Perform weighted fusion on the feature vectors corresponding to each weight.
[0109] In some embodiments, the feature fusion network is a P-PANet network, and the structure of P-PANet is as follows: Figure 6 As shown in the figure, the P-PANet network is used to progressively fuse features layer by layer. The network has a total of 17 layers. P-PANet is mainly composed of a progressive feature pyramid network (AFPN) and a path aggregation network (PAN). It promotes direct feature fusion between non-adjacent layers, thereby preventing the loss or degradation of feature information during transmission and interaction.
[0110] Extracting features from layers 4, 6, and 10 from the feature extraction network, we get a set of features of different scales represented as {P1, P2, P3}. When performing feature fusion, we first input low-level features P1 and P2 into the feature pyramid network, and then add P3. After the feature fusion step, we generate a set of multi-scale features {P4, P5, P6}. The semantic gap between features at non-adjacent levels is larger than that between features at adjacent levels, especially between the bottom and top features.
[0111] Therefore, the effect of directly fusing features of non-adjacent levels is poor, and it is unreasonable to directly use P1, P2 and P3 for feature fusion. The architecture of AFPN is progressive, which will make the semantic information of features at different levels closer in the process of progressive fusion, thereby alleviating the above problems.
[0112] The P-PANet network is equipped with an ASFF module. The structure of the ASFF module is as follows: Figure 7 As shown in the figure, in the process of multi-level feature fusion, ASFF assigns different spatial weights to features at different levels, enhances the importance of key levels, and reduces the impact of conflicting information from different targets. Its input and output are shown below:
[0113] set up Represents the feature vector from level n to level l at position (i, j). By adaptively fusing multi-layer features, the feature vector obtained is is the eigenvector and The linear combination of is defined as follows:
[0114]
[0115] in, and Respectively represent the spatial weights of the features of the previous level at level l, and satisfy the constraints
[0116] S3. Train the small object extraction model.
[0117] During the training of the small object extraction model in step S3:
[0118] The batch size of the GPU is set to 16; the decay strategy selects cosine annealing; the optimizer selects the stochastic gradient descent algorithm; the initial learning rate is set to 0.01; the dynamic parameter is set to 0.937; and the number of training rounds is set to train 350 traversals.
[0119] The early stopping strategy is used to regulate the training process of the small target extraction model. Specifically:
[0120] During the training process, the performance indicators of the small target extraction model are monitored. When the performance indicators do not improve significantly within 100 consecutive traversals, the training is terminated and completed.
[0121] S4. Use the trained small target extraction model to extract the target information to be extracted from the real-time satellite remote sensing image, and complete the precise monitoring and management of the power system based on the target information to be extracted.
[0122] Before extracting the target information to be extracted from the real-time satellite remote sensing image, the performance indicators of the small target extraction model are verified. If the verification result is up to standard, the small target extraction model is put into use. Otherwise, the small target extraction model is adjusted until the performance indicators of the small target extraction model meet the standards.
[0123] The performance indicators include:
[0124] Precision: the proportion of samples predicted as positive by the small object extraction model that are correctly predicted;
[0125] Recall rate is the proportion of samples that are actually positive and are correctly predicted by the small target extraction model;
[0126] F1 score, the harmonic mean of precision and recall;
[0127] mAP0.5, used to measure the accuracy of the model in the detection task.
[0128] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned method for extracting transmission towers from high-resolution satellite remote sensing images.
[0129] Compared with the prior art, the above method for extracting targets from high-resolution satellite remote sensing images has the following beneficial effects:
[0130] The LSKA-AFF module was constructed. The network has good feature extraction performance and the feature enhancement module has more accurate recognition ability for the features of the target to be extracted.
[0131] A P-PANet feature fusion network structure was constructed, which effectively avoided information loss or degradation during feature transmission and interaction, and realized multi-scale feature fusion of targets to be extracted under complex backgrounds.
[0132] In order to realize the extraction of targets to be extracted from high-resolution satellite remote sensing images, the above-mentioned LSKAFF-YOLO network is proposed, and its main conclusions are as follows:
[0133] In order to expand the receptive field of the model and effectively utilize the contextual information of satellite remote sensing images, the present invention constructs LSKA-AFF in the feature extraction network, which uses the attention feature fusion module (AFF) to enhance its residual connection. LSKA-AFF uses the decomposed large convolution kernel to expand the receptive field of the model and capture the long-distance dependency between global and local features, thereby achieving precise positioning of the target to be extracted in a complex background.
[0134] In order to solve the problem of semantic difference misalignment caused by simple upsampling and lateral connection in the fusion process of feature pyramid (FPN), the feature fusion network of the baseline model is replaced by P-PANet. P-PANet effectively reduces the semantic difference between features of different scales by progressively fusing finer features layer by layer, and reduces the problems of information loss and feature degradation.
[0135] The performance of LSKAFF-YOLO was evaluated on the public SRSPTD dataset and the self-built GFTD dataset. The experimental results show that mAP0.5 reached 88.8% and 94.6% respectively, and LSKAFF-YOLO significantly improved the detection accuracy of the target to be extracted in high-resolution satellite remote sensing images.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and specification of the present invention.
[0137] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for extracting transmission towers from high-resolution satellite remote sensing images, characterized in that: The steps include: Acquire satellite remote sensing images with target location information to be extracted; Constructing a small target extraction model, the small target extraction model consists of a feature extraction network, a feature fusion network, and a target detection head; wherein the feature extraction network captures feature maps from satellite remote sensing images and weights the feature maps through an attention mechanism, and the feature fusion network performs feature fusion on the weighted feature maps; Train the small object extraction model; The trained small target extraction model is used to extract the target information to be extracted from the real-time satellite remote sensing image, and the precise monitoring and management of the power system is completed based on the target information to be extracted.
2. The method for extracting transmission towers from high-resolution satellite remote sensing images according to claim 1, characterized in that: The process of obtaining the satellite remote sensing image with the target location information to be extracted is as follows: Acquire multiple high-resolution satellite remote sensing images of the target to be extracted, and combine the multiple satellite remote sensing images into a data set; The locations of the targets to be extracted in all images of the data set are marked, and the marked data set is divided into a validation set and a training set in a ratio of 1:
9.
3. The method for extracting transmission towers from high-resolution satellite remote sensing images according to claim 1, characterized in that: The feature extraction network consists of multiple convolutional layers and LSKA-AFF modules. The satellite remote sensing images processed by multiple convolutional layers are used as the original feature map, and the following operations are performed: The original feature map is sequentially passed through the batch normalization layer, 1×1 convolution, Gaussian error linear unit, LSKA module, and 1×1 convolution to obtain the first feature map; The original feature map and the first feature map are passed into the AFF module for processing to obtain a second feature map; The second feature map is sequentially passed through a batch normalization layer and a convolutional feed-forward layer to obtain a third feature map; The second feature map and the third feature map are passed into the AFF module for processing to obtain the output feature map.
4. The method for extracting transmission towers from high-resolution satellite remote sensing images according to claim 3 is characterized in that: The LSKA module receives the feature map output by the Gaussian error linear unit and performs the following operations: The feature map is convolved with the depth convolution after two layers of splitting to get the output Will The convolution operation is performed with the depth-wise separable convolution after the two layers are split, and the output A is obtained after 1×1 convolution C ; Output feature maps F and A C The element-wise product of .
5. The method for extracting transmission towers from high-resolution satellite remote sensing images according to claim 3 is characterized in that: The AFF module receives two feature maps and performs the following operations: The weight array is obtained by summing the elements of the two feature maps; Performing an element-wise product of a weight array and a feature map, and performing an element-wise product of a reverse array and another feature map, wherein the reverse array = 1-weight array; Add the results of the two element-wise multiplications to obtain the weighted feature map.
6. The method for extracting transmission towers from high-resolution satellite remote sensing images according to claim 1, characterized in that: The feature fusion network consists of a progressive feature pyramid network and a path aggregation network. The working process of the feature fusion network is as follows: Get features from multiple convolutional layers of the feature extraction network; Input low-level features into the front progressive feature pyramid network; The output of the front progressive feature pyramid network and the high-level features are input into the rear progressive feature pyramid network to obtain multi-scale features; The multi-scale features are input into the path aggregation network to obtain the detection results of the target to be extracted.
7. The method for extracting transmission towers from high-resolution satellite remote sensing images according to claim 1, characterized in that: The feature fusion network includes an ASFF module, which receives the weighted feature map and performs the following operations: Extract feature vectors at the same position in feature maps with different weights; Re-adjust the weight corresponding to each eigenvector until the weights of all eigenvectors meet the constraint condition, where the sum of all weights is 1; Perform weighted fusion on the feature vectors corresponding to each weight.
8. The method for extracting transmission towers from high-resolution satellite remote sensing images according to claim 1, characterized in that: During the training of the small object extraction model: The attenuation strategy selects cosine annealing; the optimizer selects the stochastic gradient descent algorithm; the performance indicators of the small target extraction model are monitored, and when the performance indicators do not improve significantly within 100 consecutive traversals, the training is terminated and completed.
9. The method for extracting transmission towers from high-resolution satellite remote sensing images according to claim 1, characterized in that: Before extracting the target information to be extracted from the real-time satellite remote sensing image, the performance indicators of the small target extraction model are verified. If the verification result is up to standard, the small target extraction model is put into use. Otherwise, the small target extraction model is adjusted until the performance indicators of the small target extraction model meet the standards.
10. The method for extracting transmission towers from high-resolution satellite remote sensing images according to claim 9, characterized in that: The performance indicators include: Precision, the proportion of samples predicted as positive by the small target extraction model that are correctly predicted; recall, the proportion of samples that are actually positive that are correctly predicted by the small target extraction model; F1 score, the harmonic mean of precision and recall; mAP0.5, used to measure the accuracy of the model in the detection task.
Citation Information
Patent Citations
Satellite remote sensing image small target detection method based on high-resolution characteristic self-attention
CN117036980A
Electric power tower remote sensing target detection method based on big kernel selection feature fusion network
CN118212546A
Method and system for detecting bird species caused by mistaken collision with power transmission line based on deep learning
CN118247554A
Road damage detection method based on improved YOLOv8
CN118521869A
Mask R-CNN-based high-frequency detail enhancement remote sensing image feature extraction improvement method
CN118982677A